Qwen 3.8 Max scores below Kimi K3 on the intelligence index yet costs a third more per task
Benchmark scores do not predict your bill, and Qwen 3.8 Max is the cleanest example yet.
The numbers
Artificial Analysis has now measured Qwen 3.8 Max, the open-weights model that drew attention at launch, and the ranking is the least interesting part. On the Artificial Analysis Intelligence Index at max effort settings, Claude Opus 5 scores 63, Claude Fable 5 scores 62, GPT-5.6 Sol scores 61, Kimi K3 scores 60 and Qwen 3.8 Max scores 58.
The cost column is where it gets instructive. Kimi K3 runs $0.84 per task. Qwen 3.8 Max runs $1.13. So Qwen scores two points lower than Kimi K3 and costs about a third more per task.
Why a higher score can cost more
The Decoder reports that Qwen 3.8 Max now takes 64 steps per task where it used to take 14, and that its input tokens grew roughly 15 times, because the harness resends the whole conversation history at every step. More steps and more re-sent context buy benchmark points, and the user pays for every one of them.
The pattern holds at the top end too. Claude Opus 5 leads the index at 63 but costs $2.34 per task, so the expensive-and-best trade is real. And a correction worth making: posts claiming Qwen took first place ahead of Opus 5 do not match what the data shows.
What this means for anyone shipping
Index position tells you what a model can do. Steps per task and tokens per step tell you what it will cost you to do it. Two models a couple of points apart on a leaderboard can sit far apart on an invoice, and the direction of the gap is not predictable from the score. For teams choosing a model for production, cost per task deserves the same scrutiny as the benchmark number, because the leaderboard omits the variable that shows up on the bill.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.