Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
BSC-LT releases Visual Salamandra, a 7B multimodal model that extends the Salamandra LLM to handle images and video, using SigLIP encoder and late-fusion architecture. The model is trained on multilingual data with a focus on European languages and is released under Apache 2.0 license.
From the source
The Language Technologies Lab takes a major step forward in multimodal artificial intelligence with the release of Visual Salamandra, extending the capabilities of the Salamandra large language model (LLM) to both images and video. Visual Salamandra is based on the 7 billion parameters foundational model maintaining its compactness and efficiency while extending it to multimodal tasks.
huggingface.co