Android

50 updates on Android.

  1. MediaTek launches Dimensity 9600 Pro, a 2nm chip for on-device models up to 30B parameters

    MediaTek says the NPU 1090 in the 2nm Dimensity 9600 Pro raises LLM prefill by 51% and supports on-device models with up to 30B parameters.

  2. Arm recaps Arm Create China and shows Qwen3-TTS 0.6B running on a vivo X300 CPU

    Arm's Create recap shows Qwen3-TTS 0.6B on a vivo X300 CPU with SME2 at a 1.2x real-time factor and 4.7 GB peak memory.

  3. llama.cpp Hexagon NPU backend tested on a Snapdragon 8 Gen 3 phone

    A user report puts Gemma 3 4B at 12.5 tokens per second of generation on a OnePlus 12 using llama.cpp’s Hexagon NPU backend, at about CPU speed but without the heat.

  4. Arm unveils Mali G2-Ultra NX GPU with neural accelerators in every shader core

    Arm says the Mali G2-Ultra NX runs neural upscaling and frame generation on accelerators inside its shader cores, with up to 24% higher benchmark performance.

  5. Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models

    Arm says its C2-Ultra CPU with two SME2 units delivers up to 1.7x performance across the latest AI models and finishes an agentic workflow 24% faster.

  6. Google lists first Gemini Nano 4 phones, requires Nano 3 for Gemini Intelligence

    ML Kit GenAI documentation now names nano-v4 devices and sets Nano v3 or greater, plus 12 GB of RAM, as the requirement for Gemini Intelligence.

  7. Ornith releases Ornith-1.5, a 9B model with a mobile build for iPhone and Android

    Ornith-1.5 comes in 9B, 35B and 397B sizes, and the 9B model scores 70.6 on SWE-bench Verified and has a mobile build for iPhone and Android.

  8. Online-SDFT reports 70.28% routing accuracy and updates a LoRA adapter on the phone

    I-Ju Lin and Zhang-Wei Hong measure 70.28% accuracy against 52.78% for the best baseline, and run the LoRA update itself on an Android phone.

  9. RikkaHub Agent test: Android phone agent compiles whisper.cpp on its own

    XDA runs RikkaHub Agent on an Oppo Find N5 against a self-hosted Qwen 3.6 27B. The agent installed dependencies and built whisper.cpp in Termux in seven minutes.

  10. Gemini Nano 4 ships on Samsung foldables with ML Kit Prompt API access

    Google says Samsung’s new foldables carry Gemini Nano 4 with support for more than 140 languages, reachable from apps through ML Kit’s Prompt API.

  11. FBLayout fine-tunes transformers on phone GPUs 2.2 to 5.7 times faster

    A MobiSys 2026 paper fine-tunes seven transformer models on phone GPUs 2.2 to 5.7 times faster than MNN, TFLite and TVM, with 4.2 times fewer cache misses.

  12. MLPerf Mobile v6.0 adds Llama tests in 1B, 3B and 8B sizes

    MLCommons added Llama tests in 1B, 3B and 8B sizes to its mobile benchmark app, reporting token throughput next to the existing vision and image tests.