Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Sentence Transformers v5.4 adds multimodal embedding and reranker support, enabling encoding and comparison of text, images, audio, and video with the same API. The update allows cross-modal similarity and retrieval, demonstrated with Qwen3-VL-Embedding-2B.
From the source
With the v5.4 update, you can now encode and compare texts, images, audio, and videos using the same familiar API.
huggingface.co