From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
New reference-free benchmark evaluates video captions using multiple-choice questions
Apple researchers introduced CapQuiz, a reference-free benchmark for evaluating video caption quality in Visual Large Language Models, along with the CapF1 composite metric.
The benchmark uses human-verified multiple-choice questions across 10 question types and 24 video domains, and experiments show it correlates better with human judgments than existing metrics.
From the source
We introduce CapQuiz, a novel reference-free benchmark that assesses captions based on their utility in answering human-verified, fine-grained, multiple-choice questions derived from the video. ... We further formulate CapF1, a composite metric that synthesizes CapP (measuring factuality) and CapR (measuring coverage). Extensive experiments demonstrate that CapQuiz correlates significantly better with human judgments than existing metrics and offers interpretable insights into model performance.
machinelearning.apple.com
Reported by