# Hugging Face — Visual Salamandra: Pushing the Boundaries of Multimodal Understanding

- Company: Hugging Face (huggingface.co)
- Announced: 2025-04-11T14:21:56+00:00
- Category: new-model
- Subject: Platform
- Open weights: yes
- Models affected: Visual Salamandra, Salamandra, Salamandra Instructed 7B
- License: Apache License, Version 2.0
- Source: https://huggingface.co/blog/BSC-LT/visualsalamandra7b
- Record: https://forck.live/items/1715-visual-salamandra-pushing-the-boundaries-of-multimodal-understanding

BSC-LT releases Visual Salamandra, a 7B multimodal model that extends the Salamandra LLM to handle images and video, using SigLIP encoder and late-fusion architecture. The model is trained on multilingual data with a focus on European languages and is released under Apache 2.0 license.

## Evidence

Verbatim from https://huggingface.co/blog/BSC-LT/visualsalamandra7b:

> The Language Technologies Lab takes a major step forward in multimodal artificial intelligence with the release of Visual Salamandra, extending the capabilities of the Salamandra large language model (LLM) to both images and video. Visual Salamandra is based on the 7 billion parameters foundational model maintaining its compactness and efficiency while extending it to multimodal tasks.

---

Record: https://forck.live/items/1715-visual-salamandra-pushing-the-boundaries-of-multimodal-understanding
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
