Benchmarks

52 updates on Benchmarks.

  1. EXO Labs benchmarks put Llama 3.1 8B at 14 tok/s on an iPhone 15 Pro

    EXO Labs published automated inference benchmarks from real devices, covering iPhone 15 Pro, Galaxy S24 Ultra, Mac mini M4 Pro clusters and desktop GPUs.

  2. BlueLM-V-3B runs a multimodal model on a Dimensity 9300 NPU in 2.2 GB

    CUHK and vivo AI Lab report 24.4 tok/s and 2.2 GB peak memory for a 3B vision-language model on a MediaTek Dimensity 9300.

  3. Hugging Face trains SmolLM2 at 135M, 360M and 1.7B on up to 11T tokens

    Hugging Face released SmolLM2 in three sizes trained on up to 11 trillion tokens, with 4-bit builds from 118 MB for on-device runtimes.

  4. Apple team finds H100 last on tokens per dollar for models up to 2B

    Seven Apple authors measured tokens per dollar for LLaMA-style models from 100M to 2B and found H100s last at every size, behind cheaper A100s.

  5. Mistral puts Ministral 3B and 8B on devices with 128k context

    Ministral 3B and 8B handle up to 128k tokens for on-device work, but only the 8B Instruct weights were published, and for research use.

  6. PalmBench finds iPhones running local LLMs about three times faster than Android phones

    A benchmark of quantised LLMs on eight phones and boards reports throughput, memory, power and heat, with iPhones ahead and 4-bit drawing more power than 3-bit.

  7. Gemma 2 2B is distilled from a larger model and scores 1126 on Chatbot Arena

    Google released a 2.6B-parameter Gemma 2 trained by distilling a larger model, and reported an Elo of 1126 on the LMSYS Chatbot Arena.

  8. Spectra ships ternary 3.9B models that match 4-bit quantisation at half the bits

    Nolano AI trained 54 models from 99M to 3.9B parameters on the same 300B tokens in ternary, quantised and half-precision form to compare them by bit size.

  9. Hugging Face releases SmolLM at 135M, 360M and 1.7B parameters

    Three base models trained on the newly released SmolLM-Corpus, with published memory footprints from 109.78 MB to 3422.76 MB.

  10. Apple's MobileCLIP-S0 encodes an image in 1.5 ms on an iPhone 12 Pro Max

    Apple timed its image-text models on an iPhone and released four variants, the weights and the reinforced DataCompDR dataset.

  11. TensorOpera releases Fox-1, a 1.6B model trained on 3 trillion tokens

    TensorOpera published a 1.6B model under Apache 2.0 and reports it ahead of Gemma-2B and Qwen1.5-1.8B on a six-benchmark average.

  12. BUPT measures 22 LLMs on four Android phones at about 200 ms per token

    A BUPT team ran 22 models from 0.5B to 7B on four Android phones with llama.cpp and found 7B models take about 4 GB and 200 ms per token.