Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face releases Docmatix, a large-scale dataset for Document Visual Question Answering (DocVQA), 240 times larger than previous datasets. It is generated from PDFA OCR dataset using Phi-3-small model for Q/A pair creation, with filtering for hallucinations. The dataset contains 2.4 million images and 9.5 million Q/A pairs from 1.3 million PDFs. Ablation studies using Florence-2 model show a 20% improvement in DocVQA performance.
From the source
With this blog we are releasing Docmatix - a huge dataset for Document Visual Question Answering (DocVQA) that is 100s of times larger than previously available.
huggingface.co