Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Google Research introduces knowledge profiling, a behavioral framework that distinguishes encoding failures from recall failures in LLMs. Using the WikiProfile benchmark of 2,150 Wikipedia-derived facts, they evaluate frontier LLMs (Gemini3, GPT-5) and find that these models encode nearly all facts but struggle to recall many, indicating that recall is the bottleneck for parametric factuality rather than encoding. The framework classifies facts into five knowledge profiles and uses encoding, knowledge, and recall measures to diagnose factual errors more precisely than standard accuracy metrics. The paper also describes the automated pipeline for constructing WikiProfile and the evaluation of 13 LLMs with and without thinking, producing approximately 4.5 million responses.
From the source
frontier LLMs encode nearly all facts, yet struggle to recall many of them.
research.google