Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Moonshot AI releases WorldVQA, a benchmark of 3,500 image-question pairs to evaluate factual visual world knowledge in multimodal LLMs, with a focus on head vs. tail distribution. Experiments show frontier models often fall below 50% accuracy, and the dataset and evaluation scripts are open-sourced.
From the source
We are releasing WorldVQA , a new benchmark designed to measure the factual correctness of Multimodal Large Language Models (MLLMs).
kimi.ai