, 1 min read

How we got first audio under 200 milliseconds

The boring engineering behind the moment a voice starts talking before you notice you asked.

Nobody waits politely for a voice. If the first sound takes longer than a breath, people repeat themselves, talk over the reply, or hang up. So we treat time to first audio as the only number that matters.

Start speaking before you finish thinking

The old way renders a whole sentence, then sends it. We split text into short phrases as it arrives and start synthesis on the first one while the rest is still being written by your model. The listener hears a voice almost immediately, and the rest catches up behind it.

Keep the models warm

Cold starts are the silent killer. Every region keeps a pool of loaded voices ready, and popular custom voices get pinned so the first request of the morning is as fast as the thousandth.

Ship bytes, not files

Audio leaves our servers in small chunks over a single connection. Your player starts on the first chunk instead of waiting for a finished file.

Step Before Now
Text to first phrase 240 ms 60 ms
Synthesis of first phrase 310 ms 95 ms
Network to first byte 120 ms 35 ms

None of this is magic. It is a lot of small things, measured every day, and a team that gets genuinely upset when a graph goes the wrong way.

Keep reading