What shipped, and when

Every release to the API, the voices and the dashboard, newest first. Faster, clearer and occasionally just less annoying.

v2.5

, API

Webhooks for long renders

Send an audiobook chapter, close the tab and get a webhook the moment the audio is ready.

Long renders no longer need an open connection. Queue the job, and we call you back when it finishes.

What you get:

  • Signed webhooks: every call carries a signature header so you can verify it came from us.
  • Retries with backoff: if your endpoint is down, we retry for up to 24 hours.
  • Progress events: optional render.progress events report how far a long job has got.
  • Replay from the dashboard: resend any event from the last 30 days with one click.

Webhooks work with every plan. Set up your first endpoint under Settings, then Webhooks.

Read release

v2.4

, Voices

Twenty new languages

Every stock voice now speaks Polish, Finnish, Greek, Thai and sixteen more, with native accents.

The stock library now covers 120 languages and accents. Every existing voice picks up the new ones, so you keep the same character across markets.

Highlights:

  • Native accents: each language was tuned with native speakers, not translated from English prosody.
  • Mixed sentences: switch languages mid-sentence and the voice follows, handy for names and product terms.
  • Auto detect: leave language empty and we pick it from the text.
  • Dictionary support: pronunciation dictionaries work in every new language from day one.

Custom voices from Voice Lab get the new languages too. Expect the best results in languages your sample already sounds close to.

Read release

v2.3

, Billing

Usage alerts and spend caps

Get warned before a busy week turns into a surprise invoice, and set a hard cap per project.

You can now see where your characters go and stop them before they cost more than you planned.

What’s new:

  • Usage alerts: get an email at 50, 80 and 100 percent of your monthly allowance, or pick your own thresholds.
  • Spend caps: set a hard monthly limit per project. Requests over the cap return a clear 402 instead of quietly billing.
  • Per key breakdown: the usage page now splits characters by API key, so you can spot the noisy one.
  • CSV export: download daily usage for any range and drop it straight into your finance sheet.

Alerts are on by default for every workspace owner. Caps are off until you set one.

Read release

v2.2

, Voices

Pronunciation dictionaries

Teach every voice how to say your brand names, products and tricky words once, then forget about it.

Pronunciation dictionaries are here. Add a word, write how it should sound, and every voice in your project uses it.

What you can do:

  • Plain spelling: write words the way they sound, no phonetic alphabet needed.
  • Per project: attach different dictionaries to different projects or languages.
  • Import and export: keep dictionaries as plain files next to your code.
  • Instant updates: changes apply to the next request, no retraining.

Start with your company name, your product names and the surnames of your biggest customers. It takes ten minutes and fixes the mistakes people notice most.

Read release

v2.1

, Voice Lab

Voice Lab: clone a voice from thirty seconds

Create a custom voice from a short recording, with spoken consent checks and per-workspace ownership built in.

Voice Lab turns a thirty second recording into a custom voice you can use anywhere in the API.

Key highlights:

  • Short samples: thirty seconds of clean speech is enough. Two minutes gets you noticeably closer.
  • Spoken consent: the voice owner reads a random sentence aloud, and we match it to the sample before cloning.
  • Private by default: custom voices belong to one workspace and can be deleted at any time.
  • Watermarked output: every clip carries an inaudible mark that identifies where it came from.

Voice Lab is available on the Pro, Studio and Blackbox plans. Free workspaces can try it with one voice.

Read release

v2.0

, Streaming

Streaming v2: first audio 40% faster

A rebuilt streaming pipeline that starts speaking sooner, holds intonation across chunks and reconnects on its own.

Streaming v2 is live for every workspace. Nothing to migrate: existing WebSocket connections pick it up automatically.

What changed:

  • Faster first audio: median time to first byte dropped from 290 ms to 170 ms.
  • Smoother joins: phrases streamed separately now share intonation, so long answers no longer sound stitched together.
  • Automatic reconnects: dropped connections resume mid-turn without repeating what was already spoken.
  • Flush on demand: a new flush message finishes the current turn immediately, handy for interruptions.

If you were padding text to get better intonation, you can stop. Send it as it arrives.

Read release