# Kyutai — MoshiVis: Teaching Moshi to Converse about Images

- Company: Kyutai (kyutai.org)
- Announced: 2025-03-21
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://kyutai.org/moshivis
- Record: https://forck.live/items/18357-moshivis-teaching-moshi-to-converse-about-images
- Subject: Moshi / Pocket TTS

Amélie Royer*, Moritz Böhle*, Gabriel de Marmiesse, Laurent Mazaré, Neil Zeghidour, Alexandre Défossez, and Patrick Pérez. With Moshi , we paved a new road towards natural speech-to-speech interactions with large language models. Today, we introduce MoshiVis , an open-source Vision Speech Model (VSM) with the same low-latency and natural conversation skills as Moshi, with the additional ability to discuss visual inputs. To do so, MoshiVis augments Moshi with lightweight adaptation modules for visual inputs which are trained with a data-efficient one-stage training pipeline. Here, we provide additional details and qualitative samples to accompany our preprint on MoshiVis. there! How is it going? · · · · I see two green metal structures with a mesh top, and they're surrounded by large trees. · · · the background, you can see a building with a light brown exterior and a black roof, which appears to be made of stone. · · · Based on the large trees and the warm colors of the building, I'd say we're in spring or early summer. · · · · · · The structures are green metal cages with a mesh top, likely for climbing or play. …

---

Record: https://forck.live/items/18357-moshivis-teaching-moshi-to-converse-about-images
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
