From the source
When we launched Gradium in December 2025, we started with the core models that power voice agents: best-in-class Text-to-Speech (TTS) and Speech-to-Text (STT).
But having great STT and TTS models is only half the battle.
To actually build a voice agent, you need to orchestrate these components: wire the transport layer, handle interruptions, manage conversation flow, coordinate tool calls, and sync everything in real time.
In order to spin up client demos quickly, we built a minimal internal framework to handle the orchestration and design custom voice demos easily.
Today we're sharing it with the community.
Gradbot is an open-source framework for prototyping voice agents in minutes.
Whether you're building a 3D NPC game or a travel booking assistant , Gradbot lets you go from idea to working voice experience in around 50 lines of code.
Gradbot Core Voice agents need three streams running in sync: listening to the user (speech-to-text), deciding what to say (LLM inference), and speaking back (text-to-speech).
At the center of Gradbot is a multiplexing engine written in Rust that coordinates these three pipelines in real time.
…





