Chips
11 updates on Chips.
MediaTek launches Dimensity 9600 Pro, a 2nm chip for on-device models up to 30B parameters
MediaTek says the NPU 1090 in the 2nm Dimensity 9600 Pro raises LLM prefill by 51% and supports on-device models with up to 30B parameters.
Arm recaps Arm Create China and shows Qwen3-TTS 0.6B running on a vivo X300 CPU
Arm's Create recap shows Qwen3-TTS 0.6B on a vivo X300 CPU with SME2 at a 1.2x real-time factor and 4.7 GB peak memory.
iPhone 18 Pro: A20 Pro adds a dual 16-core Neural Engine
Apple says the A20 Pro carries 32 Neural Engine cores in total, double the AI processing power of A19 Pro, with 50 percent more memory bandwidth.
Arm unveils Mali G2-Ultra NX GPU with neural accelerators in every shader core
Arm says the Mali G2-Ultra NX runs neural upscaling and frame generation on accelerators inside its shader cores, with up to 24% higher benchmark performance.
Arm unveils CSS for Mobile 2 with C2 CPU cluster, up to 1.7x faster on AI models
Arm says its C2-Ultra CPU with two SME2 units delivers up to 1.7x performance across the latest AI models and finishes an agentic workflow 24% faster.
Pixel 11 series: Tensor G6 adds 50 percent more TPU compute
Google says Tensor G6 with the latest Gemini Nano model processes on-device AI tasks up to 3.5 times faster while using up to 3.5 times less energy.
A19 Pro puts Neural Accelerators in every GPU core
Apple says the iPhone 17 Pro chip pairs Neural Accelerators in each of six GPU cores with a 16-core Neural Engine to run large local language models.
ROMA keeps a 4-bit 3B model in on-chip ROM and reports 31,800 tok/s in synthesis
A synthesised 7 nm accelerator holds a 4-bit 3B Llama in read-only memory and the LoRA adapter in SRAM, with no chip and no FPGA prototype built.
GenAI at the edge survey lists 12 accelerators, 8 of them only simulated
A Johns Hopkins and Duke survey of generative AI on edge devices puts peak accelerator efficiency at 74.34 TOPS/W, with 8 of 12 designs only simulated.
BitNet b1.58 gives every weight three values and runs 3B in 2.22 GB
Microsoft trained models whose every weight is -1, 0 or 1, which replaces multiplication with addition, and reports parity with full-precision Llama from 3B.
Snapdragon 8 Gen 3 targets 10-billion-parameter models on device
Qualcomm says the new flagship runs generative models with up to 10 billion parameters on device and reaches up to 20 tokens per second for LLMs.