# Hugging Face — Docmatix - a huge dataset for Document Visual Question Answering

- Company: Hugging Face (huggingface.co)
- Announced: 2024-07-18T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Models affected: Florence-2, Phi-3-small, Phi-3 model, Idefics2
- Source: https://huggingface.co/blog/docmatix
- Record: https://forck.live/items/1851-docmatix-a-huge-dataset-for-document-visual-question-answering

Hugging Face releases Docmatix, a large-scale dataset for Document Visual Question Answering (DocVQA), 240 times larger than previous datasets. It is generated from PDFA OCR dataset using Phi-3-small model for Q/A pair creation, with filtering for hallucinations. The dataset contains 2.4 million images and 9.5 million Q/A pairs from 1.3 million PDFs. Ablation studies using Florence-2 model show a 20% improvement in DocVQA performance.

## Evidence

Verbatim from https://huggingface.co/blog/docmatix:

> With this blog we are releasing Docmatix - a huge dataset for Document Visual Question Answering (DocVQA) that is 100s of times larger than previously available.

---

Record: https://forck.live/items/1851-docmatix-a-huge-dataset-for-document-visual-question-answering
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
