OpenAI previews Ultrafast, running GPT-5.6 Sol at up to 750 tokens per second
750 tokens per second. OpenAI is previewing Ultrafast, a new tier that runs GPT-5.6 Sol up to 14x faster.
What was announced
The announcement came on August 13, 2026. Ultrafast is a service tier in the OpenAI API, powered by Cerebras hardware, generating up to 750 output tokens per second, up to 14 times faster than Standard processing. It is the same model with radically less waiting, and it launches first in the API. OpenAI describes this week’s release as an early look.
Why speed is the product
The obvious beneficiaries are agents. A coding agent that iterates 14x faster feels like a different product. Voice assistants stop pausing mid-sentence. Long documents get drafted before the user finishes phrasing the follow-up.
Latency has quietly become the constraint that decides what agentic products are viable. Multi-step workflows compound every generation delay, so a 14x speedup at the token level can turn a minutes-long agent run into seconds. That changes product categories, not just benchmarks.
The hardware signal
The other message is in the stack. Cerebras wafer-scale chips landing inside OpenAI’s serving infrastructure is a very public vote against GPU-only inference. Inference is diversifying away from the default answer, and speed tiers are how that diversification shows up as a product line. If Ultrafast holds up outside the preview, expect the rest of the market to need an answer to the number 750.
Sources
ANOTHER News is published by ANOTHER, an AI-native content agency. Daily coverage also runs on Instagram.