llama.cpp ships multi-token prediction for DeepSeek V4-Flash
Release b10228 landed on August 2, 2026, merging pull request #25784 titled “DeepseekV4 MTP + DSpark”. A build the same day added automatic discovery of DSpark sidecars, and b10237 extended MTP to DeepSeek V3.2 on August 3.
Why this matters if you run models yourself
Multi-token prediction lets a model draft several tokens per forward pass instead of one, so generation speeds up without buying a bigger GPU. It arrives two days after DeepSeek released V4-Flash-0731 on July 31, the version that pushed Terminal Bench 2.1 from 61.8 to 82.7 while pricing stayed at $0.14 per million input tokens and $0.28 per million output.
The gap is now measured in days
The distance between a strong open model appearing and that model running well on hardware you already own used to be measured in months. It is now days. llama.cpp, sitting past 123,000 GitHub stars, is the piece of infrastructure that keeps making that true.
Separately, community testers put V4-Flash-0731 ahead of Fable 5, Sol and Kimi K3 on a chess benchmark this weekend. Community-run, not official, so weight it accordingly.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.