From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Rubric-based reward framework improves open-domain QA quality.
Apple researchers introduce a rubric-based reward framework for open-domain question answering that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions.
The approach improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4% across three evaluation axes (composition, grounding, and instruction-following).
Conditioning rubrics on retrieved evidence improves factual support, while decomposing rubrics into quality-specific dimensions further improves coherence, organization, and adherence to query requirements.
From the source
Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and over flat rubric variants by 4%, with consistent gains across all evaluation datasets.
machinelearning.apple.com