Cartesia released Sonic-3.6, an updated voice model that improves naturalness and quality, with listeners preferring it over competitors in blind head-to-head tests across multiple languages.
Cartesia released two new features for its Ink-2 streaming STT model: keyterm prompting to improve transcription accuracy on domain-specific entities, and configurable turn detection to tune latency…
The evaluation of text-to-speech models involves multiple dimensions including correctness, audio quality, naturalness, and robustness, and metrics like word error rate saturate as models improve.
Cartesia released Ink-2, a speech-to-text model for voice agents that is ranked #1 on Artificial Analysis's streaming leaderboard for lowest word error rate.
Cartesia published a guide to voice AI terminology, explaining concepts such as STT, VAD, TTS, and turn detection, and referencing its own models Ink, Ink-2, and Sonic.
Cartesia published a guide on selecting voice AI models for enterprise voice agents, emphasizing that real-world performance depends on use-case-specific conditions such as telephony infrastructure,…
Cartesia and Goomba Lab introduced Mamba-3, a state space model designed with inference efficiency as the primary goal.
Cartesia describes its adoption of multiple coding agents and identifies the remaining bottleneck as agents' inability to close feedback loops by checking production metrics, dashboards, or…
Cartesia announced that its text-to-speech platform is now GDPR compliant, reflecting its commitment to data protection and privacy.
Cartesia introduced Line, a code-first voice agent development platform. Line is available to all developers, with subscription tiers receiving prepaid credits.
Cartesia announced a research collaboration on hierarchical networks (H-Nets), a new architecture that learns to segment and compress raw data into meaningful concepts.
Cartesia introduced Ink, a family of streaming speech-to-text models, with the debut model Ink-Whisper optimized for low-latency transcription in conversational settings.
Cartesia introduced two new features for its platform: Organizations, which enables teams to share API keys, voices, and billing under one account, and Dashboards, which provides real-time visibility…
Cartesia introduced Professional Voice Clones (PVCs) built on the Sonic text-to-speech model, available on the Startup plan and above.
Cartesia released v2.0.0 of its Python SDK, improving the developer experience for using Cartesia's AI voice capabilities with Python.
Cartesia was named to the seventh annual Enterprise Tech 30 list by Wing Venture Capital, which recognizes the most promising private enterprise tech companies across all stages of maturity.
Cartesia announced a $64 million Series A funding round led by Kleiner Perkins. The company also launched Sonic 2.0, a voice generation model built on a new state space model architecture, with lower…
Cartesia released a technical report introducing MOHAWK, a multi-stage architecture distillation method that converts Transformer models into efficient Mamba-2 variants.
Cartesia published a tutorial on building voice AI agents using its API and the Sonic model. The guide covers setting up speech-to-text, response generation, and text-to-speech with Cartesia's TTS…
Cartesia published its first State of Voice report for 2024, highlighting key infrastructure breakthroughs and emerging use cases in voice AI, and looking ahead to 2025.
Cartesia announced a $27M seed round led by Index Ventures with participation from multiple investors.
Cartesia hosted a series of hackathons in October 2024, bringing together over 2,000 builders in San Francisco and awarding $20,000 in prizes for projects built on its Sonic voice model.
Cartesia released Voice Changer, a new model that transforms the voice of any audio clip while preserving the original delivery and emotion.
Cartesia announced the alpha release of 8 new languages for Sonic Multilingual, expanding total language support to 15.
Cartesia announced Edge, an open-source library for on-device state space models; Rene, an open-source 1.3B parameter language model; and Sonic On-Device, a generative voice model in private beta, as…
Cartesia released Sonic, a low-latency voice model that generates lifelike speech with a latency of 135ms, and includes a web playground and API.
Cartesia presents Based, a recurrent architecture that uses sliding window attention and linear attention to outperform prior sub-quadratic models on recall-intensive tasks while achieving faster…
Cartesia and Together released Mamba-3B-SlimPJ, a 2.8B parameter state-space model trained on 600B tokens of the SlimPajama dataset, under an Apache 2.0 license.