# Hugging Face — Vision Language Models Explained

- Company: Hugging Face (huggingface.co)
- Announced: 2024-04-11T00:00:00+00:00
- Subject: Platform
- Source: https://huggingface.co/blog/vlms
- Record: https://forck.live/items/1910-vision-language-models-explained

This blog post provides an introduction to vision language models, including their architecture, a table of open-source models, evaluation benchmarks, and guidance on fine-tuning.

## Evidence

Verbatim from https://huggingface.co/blog/vlms:

> Vision language models are models that can learn simultaneously from images and texts to tackle many tasks, from visual question answering to image captioning.

---

Record: https://forck.live/items/1910-vision-language-models-explained
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
