# Hugging Face — Multimodal Embedding & Reranker Models with Sentence Transformers

- Company: Hugging Face (huggingface.co)
- Announced: 2026-04-09T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Models affected: Qwen3-VL-2B, Qwen/Qwen3-VL-Embedding-2B
- Source: https://huggingface.co/blog/multimodal-sentence-transformers
- Record: https://forck.live/items/1522-multimodal-embedding-reranker-models-with-sentence-transformers

Sentence Transformers v5.4 adds multimodal embedding and reranker support, enabling encoding and comparison of text, images, audio, and video with the same API. The update allows cross-modal similarity and retrieval, demonstrated with Qwen3-VL-Embedding-2B.

## Evidence

Verbatim from https://huggingface.co/blog/multimodal-sentence-transformers:

> With the v5.4 update, you can now encode and compare texts, images, audio, and videos using the same familiar API.

---

Record: https://forck.live/items/1522-multimodal-embedding-reranker-models-with-sentence-transformers
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
