From the source
A real-time reasoning model for voice agents Firas Trabelsi Yanis Miraoui Samar Khanna Xinyu Zhao Gokul Gunasekaran Emily Liu Kenan Hasanaliyev Two weeks ago, alongside Mercury 2.5, we previewed Mercury Voice.
Today it's generally available for enterprise customers.
Mercury Voice is a diffusion LLM (dLLM) tuned to power voice agents.
It reasons, calls tools, and follows long system prompts while keeping latency low enough for natural conversation.
On a test set of real customer-service prompts, Mercury Voice returns its first answer token in under 320 milliseconds (median), while also beating models like GPT-6 Luna and Gemma 4 31B on a suite of agentic and voice benchmarks.
At a glance Speed: 320 ms median (p50) time to first answer token (p95: 750 ms) on production voice prompts, 5.9x faster than GPT-6 Luna (no reasoning).
Quality: Outperforms models including Gemma 4 31B, GPT-6 Luna, GLM-5.3-Flash, and Qwen3.5-397B on a composite of agentic and conversational benchmarks.
Reasoning effort: Three settings (low, medium, high) Context: 128K tokens, with up to 50K output tokens.
Price: $0.40 per million input tokens and $1.50 per million output tokens …


