Open weights

42 updates on Open weights.

  1. Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max

    Apple timed its image-text models on an iPhone and released four variants, the weights and the reinforced DataCompDR dataset.

  2. TensorOpera releases Fox-1, a 1.6B model trained on 3 trillion tokens

    TensorOpera published a 1.6B model under Apache 2.0 and reports it ahead of Gemma-2B and Qwen1.5-1.8B on a six-benchmark average.

  3. Apple releases OpenELM at 270M to 3B, with parameters spread unevenly across layers

    Apple published four models from 270M to 3B parameters with the full training framework and code to run them through MLX on Apple silicon.

  4. Microsoft runs Phi-3-mini offline on an iPhone 14 at over 12 tokens per second

    The 3.8-billion-parameter model takes about 1.8 GB at 4-bit and scores 69 percent on MMLU, which Microsoft compares to Mixtral 8x7B and GPT-3.5.

  5. Stability AI trains Stable LM 2 1.6B on seven languages and 2 trillion tokens

    The technical report details a 1.6B model pre-trained on seven languages and measures 127 tok/s for a 4-bit build on an M2 Mac mini.

  6. MobiLlama is a fully transparent 0.5B model that runs in 770 MB on a phone

    MBZUAI published a 0.5B model that shares one feed-forward block across all layers and reports 7.02 tok/s in 770 MB on a Snapdragon 685 phone.

  7. TinyLLaVA's 3.1B model outscores 7B LLaVA-1.5 on seven of nine benchmarks

    Beihang and Tsinghua researchers report a 3.1B vision-language model that beats the 7B LLaVA-1.5 on seven of nine image benchmarks.

  8. Gemma 2B and 7B open the Gemma line, built on Gemini research

    Google released Gemma 2B and 7B with an 8192-token context, weights on Kaggle and Hugging Face under a custom Gemma licence, not an open source one.

  9. MobileVLM V2 runs 1.7B, 3B and 7B vision models on 144 image tokens

    Meituan and Zhejiang University report a 1.7B vision language model at 64.2 on six benchmarks and 51.63 tok/s on an NVIDIA Jetson Orin.

  10. TinyLlama pretrains a 1.1B model on 3 trillion tokens

    Singapore University of Technology and Design trained a 1.1B model on 3 trillion tokens with 16 A100-40G GPUs and released it under Apache 2.0.

  11. Microsoft releases Phi-2, a 2.7B model it says matches models 25 times larger

    The 2.7B base model was trained on 1.4 trillion tokens in 14 days on 96 A100 GPUs, and Microsoft says it matches models up to 25 times larger.

  12. Stability AI releases StableLM Zephyr 3B for edge devices

    Stability AI tuned a 3B chat model with direct preference optimisation for edge devices and reports an MT-Bench score of 6.64.