Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
LG AI Research presents VinQA, a multimodal document QA dataset and evaluation framework presented at CVPR 2026. VinQA focuses on generating long-form answers that strategically position grounding visual elements within text, rather than appending images. The dataset includes ~130,000 pages and 42,700 QAs across seven domains, with single-page, multi-page, multi-document, and unanswerable questions. The study introduces M-GroSE, a multimodal evaluation metric. Experiments show that fine-tuning Qwen2.5-VL-7B on VinQA significantly improves its performance, approaching commercial models like GPT-4.1 and Claude 3.5 Sonnet.
From the source
As a result, the VinQA paper was successfully presented at CVPR 2026. This study focuses on two main aspects: generation and evaluation. ... VinQA, on the other hand, is not just a model. It proposes a real-world document-based multimodal QA dataset, along with a corresponding evaluation framework. ... In the VinQA test set experiment, frontier commercial models like GPT-4.1 and Claude 3.5 Sonnet still recorded the highest scores. However, once the open-source Qwen2.5-VL-7B was fine-tuned on the VinQA training set, its performance improved significantly, coming far closer to the commercial models.
lgresearch.ai