Streaming and latency

How streaming works, what affects time to first audio and how to keep conversations feeling instant.

Updated

How streaming works

Open a WebSocket, send text as you get it and play audio chunks as they arrive. Keep one connection open for the whole conversation.

What slows things down

  • Opening a new connection per sentence. Reuse the socket.
  • Waiting for full sentences. Forward text as your model writes it.
  • Large audio buffers. Start playback on the first chunk.
  • Distant regions. Let us route you, or pin the region closest to your users.

Measuring latency

Every audio message includes a timestamp. Compare it with when you sent the text to measure time to first audio from your side, and log it alongside your own metrics.

Interruptions

Real conversations overlap. When a user starts talking, stop playback and send a flush. The voice stops cleanly and the next turn is ready right away.

Questions about streaming and latency

What latency should I expect?
Median time to first audio is around 170 milliseconds from our nearest region, plus your own network time.
Which regions do you run in?
Europe, US East, US West and Asia Pacific. Requests are routed to the closest one automatically.
What happens if the connection drops?
Streaming v2 reconnects on its own and resumes the turn without repeating what was already spoken.
Can users interrupt the voice?
Yes. Send a flush message and stop playback on your side. The next turn starts immediately.