# Apple — Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

- Company: Apple (apple.com)
- Announced: 2026-09-11
- Category: research-paper
- Coverage: 1 outlet
- Announcement: yes
- Group: covered
- Source: https://machinelearning.apple.com/research/video-caption-quality
- Record: https://forck.live/items/10300-putting-captions-to-the-test-evaluating-video-caption-quality-through-multiple
- Subject: Machine Learning Research

Apple researchers introduced CapQuiz, a reference-free benchmark for evaluating video caption quality in Visual Large Language Models, along with the CapF1 composite metric. The benchmark uses human-verified multiple-choice questions across 10 question types and 24 video domains, and experiments show it correlates better with human judgments than existing metrics.

## Evidence

Verbatim from https://machinelearning.apple.com/research/video-caption-quality:

> We introduce CapQuiz, a novel reference-free benchmark that assesses captions based on their utility in answering human-verified, fine-grained, multiple-choice questions derived from the video. ... We further formulate CapF1, a composite metric that synthesizes CapP (measuring factuality) and CapR (measuring coverage). Extensive experiments demonstrate that CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance.

## Around this story

Outlets this record can name and link:

- About Amazon — The Prime Video Shop the Show Experience Introduces New Features and Expands Across Thousands of Titles: https://press.aboutamazon.com/prime-video/2026/9/the-prime-video-shop-the-show-experience-introduces-new-features-and-expands-across-thousands-of-titles

---

Record: https://forck.live/items/10300-putting-captions-to-the-test-evaluating-video-caption-quality-through-multiple
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
