# Inception — Mercury 2: the first reasoning model fast enough to pick up the phone

- Company: Inception (inceptionlabs.ai)
- Announced: 2026-07-14
- Category: new-model
- Coverage: not counted
- Announcement: yes
- Group: models
- Source: https://www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone
- Record: https://forck.live/items/18519-mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone
- Subject: Mercury
- Models affected: Mercury 2, GPT 4.1, GPT Mini, Claude Haiku, GPT OSS 120B

Inception Labs released Mercury 2, described as the world's first reasoning diffusion language model, which decodes over 1,000 tokens per second on standard NVIDIA GPUs. The model is designed for realtime voice agent applications, reducing reasoning latency from seconds to under 300 milliseconds. It offers a reasoning_effort knob with four settings and, according to the company, beats GPT 4.1 on instruction-following benchmarks while being faster.

## Evidence

Verbatim from https://www.inceptionlabs.ai/blog/mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone:

> Mercury 2 is the world's first reasoning diffusion language model, decoding 1000+ tokens per second on standard NVIDIA GPUs. That's fast enough to run a full reasoning pass and start speaking within the latency budget of a natural conversation.

---

Record: https://forck.live/items/18519-mercury-2-the-first-reasoning-model-fast-enough-to-pick-up-the-phone
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
