Back to the ticker

Meta releases Llama 3.2 1B and 3B for phones and edge devices

Meta released Llama 3.2 on September 25, 2024, including text-only 1B and 3B models built for phones and edge hardware. Both carry a 128K token context window and are aimed at summarisation, instruction following and rewriting that run locally, with the data staying on the device.

Meta built them by structured pruning from Llama 3.1 8B, then recovered quality through knowledge distillation using logits from the 8B and 70B models during pretraining. The company reports the 3B model ahead of Gemma 2 2.6B and Phi 3.5-mini on instruction following, summarisation, prompt rewriting and tool use, and puts the 1B model on a par with Gemma. In Meta’s own table the 3B model scores 77.4 on IFEval against 61.9 for Gemma 2 2B and 59.2 for Phi-3.5-mini, and 67.0 on BFCL V2 for tool use against 27.4 and 58.4.

Benchmark table comparing Llama 3.2 1B and 3B with Gemma 2 2B IT and Phi-3.5-mini IT across general, tool use, math, reasoning, long context and multilingual tasks
Table: Meta. The company measured the Gemma and Phi results itself.

The models shipped with day-one support for Qualcomm and MediaTek silicon and run on Arm, which Meta says covers 99 percent of mobile devices. Weights are on llama.com and Hugging Face, with deployment paths through PyTorch ExecuTorch for devices and Ollama for single-node setups, and the company lists more than 25 partner platforms at launch.

  1. ElastiLM resizes a shared phone LLM per request and switches in 0.31 seconds
  2. ExecuTorch alpha runs Llama 2 7B on iPhone 15 Pro and Galaxy phones
  3. Qualcomm AI Hub opens 75 pre-optimised models and profiling on hosted Snapdragon phones