Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Hugging Face and Cerebras announce a real-time speech-to-speech pipeline using Google DeepMind's Gemma 4 VLM on Cerebras hardware for low-latency inference, with an open, modular architecture that includes Nvidia's Parakeet for speech recognition and Alibaba's Qwen3TTS for text-to-speech. The system is demonstrated in a demo and the code is available on GitHub. The collaboration aims to improve latency and naturalness in voice AI interactions.
From the source
Today, we demonstrate what becomes possible when an open, modular voice AI architecture is paired with industry-leading inference speed.
huggingface.co