# Gradium — Semantic VAD: turn detection that uses meaning, not silence

- Company: Gradium (gradium.ai)
- Announced: 2026-06-02
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://gradium.ai/blog/semantic-vad
- Record: https://forck.live/items/18317-semantic-vad-turn-detection-that-uses-meaning-not-silence
- Subject: Gradium TTS / Phonon

Picture a common interaction. The user says "I'd like to cancel my flight from Boston to..." and pauses for a second to check the date on their phone. The agent jumps in: "Got it, cancelling your flight from Boston. Where to?" The user now has to interrupt the agent's interruption to finish the sentence they started. Every voice agent built on acoustic-only turn detection has this failure mode, and it's almost always traced back to one decision the system has to make every 80 milliseconds: has the user actually stopped talking? This post covers what acoustic and semantic VAD actually do, why the distinction matters, what Gradium's STT does differently, and how to wire it into your agent loop. Acoustic VAD vs semantic VAD Acoustic VAD classifies short audio frames as speech or non-speech based on signal properties: energy, spectral shape, harmonicity. It answers "is there a voice in this 20 ms window?" and nothing beyond that. Classical implementations like WebRTC VAD and Silero VAD work this way, and they work well for the task they were designed for: gating audio, suppressing background noise, deciding when to start a transcription. …

---

Record: https://forck.live/items/18317-semantic-vad-turn-detection-that-uses-meaning-not-silence
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
