# Inception — Mercury 2 for Search: Fast enough to run a hundred times per query

- Company: Inception (inceptionlabs.ai)
- Announced: 2026-08-11
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.inceptionlabs.ai/blog/mercury-2-for-search
- Record: https://forck.live/items/18517-mercury-2-for-search-fast-enough-to-run-a-hundred-times-per-query
- Subject: Mercury

For the past two years, “search” stopped meaning a ranked list of links and started meaning an agentic pipeline. Classify, rewrite, fan out, rerank, and synthesize. Fifty to a hundred LLM calls per search query, almost all sequential. That's a brutal place to put an autoregressive model: a 400ms rewrite blocks retrieval, which blocks reranking, which blocks the first word the user sees. So teams cut the pipeline down until it's shallow enough to be fast. The industry has split into two camps: synchronous search running thin, cheap models that barely think, and deep-research agents that take thirty minutes. Nobody ships the thing in the middle: a pipeline deep enough to be smart and fast enough to be useful. Mercury 2 decodes over 1000 tokens per second on standard NVIDIA GPUs. Fast enough to run every step inside the latency budget you already have. Latency is the quality budget In most products, latency and quality are separate dials. You can make the answer better by letting the user wait. In search they're the same dial, because every step that improves makes the answer is itself an LLM call that increase latency: Query rewriting finds the documents a literal match misses. …

---

Record: https://forck.live/items/18517-mercury-2-for-search-fast-enough-to-run-a-hundred-times-per-query
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
