Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post by Hugging Face demonstrates how to fine-tune Microsoft's Florence-2 vision-language model on the DocVQA dataset, showing that after fine-tuning the Levenshtein similarity score improved from 0 to 57.0. The post includes code walkthroughs and links to a demo space and GitHub repository.
From the source
In this post, we show an example on fine-tuning Florence on DocVQA. The authors report that Florence 2 can perform visual question answering (VQA), but the released models don't include VQA capability. Let's see what we can do!
huggingface.co