# Google Research — Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

- Company: Google Research (research.google)
- Announced: 2026-08-12T09:51:00+00:00
- Category: research-paper
- Subject: Research
- Models affected: Gemini3, GPT-5, Gemini-2.5-Pro
- Source: https://research.google/blog/empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality/
- Record: https://forck.live/items/3731-empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality

Google Research introduces knowledge profiling, a behavioral framework that distinguishes encoding failures from recall failures in LLMs. Using the WikiProfile benchmark of 2,150 Wikipedia-derived facts, they evaluate frontier LLMs (Gemini3, GPT-5) and find that these models encode nearly all facts but struggle to recall many, indicating that recall is the bottleneck for parametric factuality rather than encoding. The framework classifies facts into five knowledge profiles and uses encoding, knowledge, and recall measures to diagnose factual errors more precisely than standard accuracy metrics. The paper also describes the automated pipeline for constructing WikiProfile and the evaluation of 13 LLMs with and without thinking, producing approximately 4.5 million responses.

## Evidence

Verbatim from https://research.google/blog/empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality/:

> frontier LLMs encode nearly all facts, yet struggle to recall many of them.

---

Record: https://forck.live/items/3731-empty-shelves-or-lost-keys-recall-is-the-bottleneck-for-parametric-factuality
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
