Open weights
42 updates on Open weights.
Alibaba gives Qwen3's 0.6B and 1.7B models a reasoning switch
Qwen3-0.6B and Qwen3-1.7B carry the family's switch between a reasoning mode and a fast mode, with 32K context and Apache 2.0 weights.
HuggingSnap describes what the iPhone camera sees with a 500M model on the phone
Hugging Face released an iPhone app that runs SmolVLM2 at 500M parameters through MLX, describing camera scenes, photos and video with no cloud call.
Gemma 3 adds a 1B size that fits in 0.5 GB as an int4 checkpoint
Google released Gemma 3 at 1B, 4B, 12B and 27B with quantisation-aware int4 checkpoints of 0.5 GB and 2.6 GB for the two smallest sizes.
PhoneLM searches for a fast architecture before training it and hits 58 tok/s
BUPT researchers picked their 0.5B and 1.5B transformer shapes by measuring speed on a Snapdragon 8 Gen 3 first, then pre-training the winner.
Hugging Face trains SmolLM2 at 135M, 360M and 1.7B on up to 11T tokens
Hugging Face released SmolLM2 in three sizes trained on up to 11 trillion tokens, with 4-bit builds from 118 MB for on-device runtimes.
AMD trains its first small language model, AMD-Llama-135M, on MI250 accelerators
AMD trained a 135M model from scratch on Instinct MI250 accelerators and reports up to 3.88x faster CodeLlama-7b inference when it drafts tokens.
Meta releases Llama 3.2 1B and 3B for phones and edge devices
The two lightweight models carry a 128K context window, were pruned and distilled from Llama 3.1, and shipped with day-one Qualcomm and MediaTek support.
Ai2 releases OLMoE, 7B parameters with 1B active per token
Ai2's mixture-of-experts model holds 6.9B parameters but runs 1.3B per token, and its iOS app runs it offline on an iPhone 15 Pro or newer.
Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena
Google released a 2.6B-parameter Gemma 2 trained by distilling a larger model, and reported an Elo of 1126 on the LMSYS Chatbot Arena.
Spectra ships ternary 3.9B models that match 4-bit quantisation at half the bits
Nolano AI trained 54 models from 99M to 3.9B parameters on the same 300B tokens in ternary, quantised and half-precision form to compare them by bit size.
Hugging Face releases SmolLM at 135M, 360M and 1.7B parameters
Three base models trained on the newly released SmolLM-Corpus, with published memory footprints from 109.78 MB to 3422.76 MB.
Alibaba builds Qwen2's 0.5B and 1.5B sizes for phones, earphones and glasses
Alibaba built Qwen2-0.5B and Qwen2-1.5B for smartphones, earphones and smart glasses, with 32K context and Apache 2.0 weights.