Qualcomm
19 updates on Qualcomm.
PowerInfer-2 runs a 47B model on a OnePlus 12 at 11.68 tokens per second
Shanghai Jiao Tong University researchers report a 47B model decoding at 11.68 tokens per second on a OnePlus 12, with weights streamed from flash.
ExecuTorch alpha runs Llama 2 7B on iPhone 15 Pro and Galaxy phones
PyTorch's edge runtime brought 4-bit Llama 2 7B to iPhone and Galaxy handsets, added early Llama 3 8B support and leaned on Apple, Arm and Qualcomm.
Qualcomm AI Hub opens 75 pre-optimised models and profiling on hosted Snapdragon phones
Qualcomm opened a library of more than 75 models tuned for Snapdragon, with compilation and profiling on real phones in its cloud and two 7B chat models listed.
Snapdragon 8 Gen 3 targets 10-billion-parameter models on device
Qualcomm says the new flagship runs generative models with up to 10 billion parameters on device and reaches up to 20 tokens per second for LLMs.
Qualcomm argues for hybrid AI, citing 10 times the cost per generative AI search query
Qualcomm argues that cloud-only inference cannot scale, and puts models of 1B to 10B parameters on phones and laptops at INT4.
Qualcomm runs Stable Diffusion on an Android phone for the first time
Qualcomm AI Research generated 512x512 images in under 15 seconds on a Snapdragon 8 Gen 2 phone after quantising the model from FP32 to INT8.
George Hotz starts tinygrad, a framework that ports to an accelerator in about 25 ops
The tiny corp framework caps its repository at 26,500 lines, ships Metal, Adreno and WebGPU backends, and runs openpilot on a Snapdragon 845 GPU.