, 1 min read

Streaming speech over WebSockets, step by step

A plain walkthrough of the streaming API, from opening a socket to playing the first chunk.

Batch requests are fine for voiceovers. For anything conversational you want streaming, so audio starts playing while text is still arriving.

Open one connection

Open a WebSocket to the streaming endpoint with your API key and a voice id. Keep it open for the whole conversation instead of reconnecting per sentence.

Send text as you get it

Push text in small pieces, a phrase or a sentence at a time. If your language model streams tokens, forward them as they come. We buffer just enough to keep intonation natural.

Play chunks immediately

Each message back contains a short piece of audio. Append it to your player’s buffer and start playback on the first one.

Close cleanly

Send a flush message at the end of a turn so the last phrase is spoken, then either keep the socket for the next turn or close it.

Message Direction Purpose
text You to us Text to speak
flush You to us Finish the current turn
audio Us to you A chunk of audio
done Us to you Turn finished

That is the whole protocol. Most teams get their first streamed sentence playing in an afternoon.

Keep reading