From the source
Today, we’re thrilled to announce that Mercury 2 is available on Azure AI Foundry , the world’s fastest reasoning language model, built to make production AI feel instant.
This release combines Inception’s breakthrough diffusion architecture with Azure’s enterprise-ready infrastructure, giving developers access to a model that delivers both exceptional speed and quality.
Built by the team behind foundational AI technologies including Flash Attention, Direct Preference Optimization, and the original diffusion models for images, Mercury represents a fundamental shift in how language models generate text.
Developers on Azure AI Foundry now have access to a model that excels in latency-sensitive applications where the user experience is non-negotiable.
A new foundation: Diffusion for real-time reasoning Mercury 2 doesn’t decode sequentially.
It generates responses through parallel refinement, producing multiple tokens simultaneously and converging over a small number of steps.
Less typewriter, more editor revising a full draft at once.
The result: >5x faster generation with a fundamentally different speed curve.
That speed advantage also changes the reasoning trade-off.
…




