# Inception — Introducing Mercury Voice

- Company: Inception (inceptionlabs.ai)
- Announced: 2026-09-29
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://www.inceptionlabs.ai/blog/introducing-mercury-voice
- Record: https://forck.live/items/18515-introducing-mercury-voice
- Subject: Mercury
- Models affected: Mercury Voice, GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash, Qwen3.5-397B, GPT-OSS-120B, GPT-4.1, Gemini 3.5 Flash-Lite
- Context window: 128K tokens
- Pricing: $0.40 per million input tokens and $1.50 per million output tokens. At launch, it's 50% off: $0.20 per million input and $0.75 per million output.

Inception made Mercury Voice generally available for enterprise customers. The diffusion LLM is tuned for voice agents, achieving a median time to first answer token of 320 ms on production voice prompts while outperforming several models on agentic and voice benchmarks. It supports 128K token context, up to 50K output tokens, and three reasoning effort settings.

## Evidence

Verbatim from https://www.inceptionlabs.ai/blog/introducing-mercury-voice:

> Mercury Voice is a diffusion LLM (dLLM) tuned to power voice agents. It reasons, calls tools, and follows long system prompts while keeping latency low enough for natural conversation. On a test set of real customer-service prompts, Mercury Voice returns its first answer token in under 320 milliseconds (median), while also beating models like GPT-6 Luna and Gemma 4 31B on a suite of agentic and voice benchmarks.

---

Record: https://forck.live/items/18515-introducing-mercury-voice
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
