# LG AI Research — [CVPR 2026] Beyond Text: How VinQA Incorporates Visual Elements into Responses for Multimodal Document QA

- Company: LG AI Research (lgresearch.ai)
- Announced: 2026-07-15T00:00:00+00:00
- Category: research-paper
- Subject: EXAONE
- Models affected: GPT-4.1, Claude 3.5 Sonnet, Qwen2.5-VL-7B
- Source: https://www.lgresearch.ai/blog/view?seq=661
- Record: https://forck.live/items/4742-cvpr-2026-beyond-text-how-vinqa-incorporates-visual-elements-into-responses

LG AI Research presents VinQA, a multimodal document QA dataset and evaluation framework presented at CVPR 2026. VinQA focuses on generating long-form answers that strategically position grounding visual elements within text, rather than appending images. The dataset includes ~130,000 pages and 42,700 QAs across seven domains, with single-page, multi-page, multi-document, and unanswerable questions. The study introduces M-GroSE, a multimodal evaluation metric. Experiments show that fine-tuning Qwen2.5-VL-7B on VinQA significantly improves its performance, approaching commercial models like GPT-4.1 and Claude 3.5 Sonnet.

## Evidence

Verbatim from https://www.lgresearch.ai/blog/view?seq=661:

> As a result, the VinQA paper was successfully presented at CVPR 2026. This study focuses on two main aspects: generation and evaluation. ... VinQA, on the other hand, is not just a model. It proposes a real-world document-based multimodal QA dataset, along with a corresponding evaluation framework. ... In the VinQA test set experiment, frontier commercial models like GPT-4.1 and Claude 3.5 Sonnet still recorded the highest scores. However, once the open-source Qwen2.5-VL-7B was fine-tuned on the VinQA training set, its performance improved significantly, coming far closer to the commercial models.

---

Record: https://forck.live/items/4742-cvpr-2026-beyond-text-how-vinqa-incorporates-visual-elements-into-responses
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
