# Hugging Face — Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models

- Company: Hugging Face (huggingface.co)
- Announced: 2024-06-24T00:00:00+00:00
- Subject: Platform
- Models affected: Florence-2
- Source: https://huggingface.co/blog/finetune-florence2
- Record: https://forck.live/items/1866-fine-tuning-florence-2-microsoft-s-cutting-edge-vision-language-models

This blog post by Hugging Face demonstrates how to fine-tune Microsoft's Florence-2 vision-language model on the DocVQA dataset, showing that after fine-tuning the Levenshtein similarity score improved from 0 to 57.0. The post includes code walkthroughs and links to a demo space and GitHub repository.

## Evidence

Verbatim from https://huggingface.co/blog/finetune-florence2:

> In this post, we show an example on fine-tuning Florence on DocVQA. The authors report that Florence 2 can perform visual question answering (VQA), but the released models don't include VQA capability. Let's see what we can do!

---

Record: https://forck.live/items/1866-fine-tuning-florence-2-microsoft-s-cutting-edge-vision-language-models
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
