From the source
For the past two years, every frontier lab has hillclimbed intelligence in the same way: letting models think for longer .
Reasoning models now burn thousands of thinking (pontificating, cerebrating, lollygagging) tokens before saying a word.
Under the dominant autoregressive decoding paradigm, every one of those tokens is generated sequentially, each requiring a full forward pass.
On the phone with a voice agent, this generation (decode) time stacks up to the point where you’d rather just hang up.
The result is a strange split in the industry.
The reasoning intelligence frontier has sprinted ahead, while realtime intelligence has stood still .
Voice is one of the only verticals where GPT 5.x, Claude Sonnet, and Gemini Pro simply fail to meet the bar because nobody wants to wait 3 seconds every time they answer a question about their dentist appointment.
Mercury 2 is the world’s first reasoning diffusion language model , decoding 1000+ tokens per second on standard NVIDIA GPUs.
That’s fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation.
We reduce the cost of reasoning from 3 seconds of dead air to only 300 milliseconds.
…




