# Moonshot AI — WorldVQA

- Company: Moonshot AI (moonshot.ai)
- Announced: 2026-02-03T00:00:00+00:00
- Category: research-paper
- Subject: Kimi
- Models affected: Kimi-K2.5
- Source: https://www.kimi.ai/blog/worldvqa
- Record: https://forck.live/items/4625-worldvqa

Moonshot AI releases WorldVQA, a benchmark of 3,500 image-question pairs to evaluate factual visual world knowledge in multimodal LLMs, with a focus on head vs. tail distribution. Experiments show frontier models often fall below 50% accuracy, and the dataset and evaluation scripts are open-sourced.

## Evidence

Verbatim from https://www.kimi.ai/blog/worldvqa:

> We are releasing WorldVQA , a new benchmark designed to measure the factual correctness of Multimodal Large Language Models (MLLMs).

---

Record: https://forck.live/items/4625-worldvqa
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
