From the source
Most voice APIs come either with a library of predefined voice presets or the ability to create a custom voice from an audio sample.
With Voice Design , you can now create natural, realistic voices from a text description (or "prompt").
Voice characteristics become output your code computes, from the same context it already uses to decide what the agent says.
Reading that context has also become inexpensive.
A new class of decision models, such as Jev from TypeSafe AI, classifies a message against labels you define for about $0.00002 per request (in the following demo), with no training.
This opens use cases that a fixed catalog of voices cannot serve well: a game that gives every procedurally generated character a voice of its own, an interactive story whose narrator changes with the scene, a support agent that answers inquiries with an appropriate tone.
Consider this latter case.
You are building an automated voice agent for the customer support line.
Most requests are routine, like asking for an address change.
Now one ticket arrives: "I have been charged twice and nobody has answered my emails."
…






