# Hugging Face — Vision Language Models (Better, faster, stronger)

- Company: Hugging Face (huggingface.co)
- Announced: 2025-05-12T00:00:00+00:00
- Subject: Platform
- Models affected: Qwen 2.5 Omni, MiniCPM-o 2.6, Janus-Pro-7B, QVQ-72B-preview, Kimi-VL-A3B-Thinking, Kimi-VL-A3B-Instruct, SmolVLM, SmolVLM2, gemma3-4b-it
- Context window: 128k token context window
- Source: https://huggingface.co/blog/vlms-2025
- Record: https://forck.live/items/1700-vision-language-models-better-faster-stronger

This blog post surveys recent developments in Vision Language Models (VLMs) over the past year, covering new model trends such as any-to-any models, reasoning models, small yet capable models, and specialized capabilities. It highlights specific models including Qwen 2.5 Omni, MiniCPM-o 2.6, Janus-Pro-7B, QVQ-72B-preview, Kimi-VL-A3B-Thinking, SmolVLM, SmolVLM2, and gemma3-4b-it, and discusses new benchmarks like MMT-Bench and MMMU-Pro.

## Evidence

Verbatim from https://huggingface.co/blog/vlms-2025:

> In this blog post, we’ll take a look back and unpack everything that happened with vision language models the past year. You’ll discover key changes, emerging trends, and notable developments.

---

Record: https://forck.live/items/1700-vision-language-models-better-faster-stronger
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
