Text-to-Speech (Bytes)
Stream audio from a complete transcript
Authorizations
Cartesia API key (sk_car_...). Get one at play.cartesia.ai/keys.
Headers
API version header.
2026-08-14 "2026-08-14"
Body
The ID of the voice.
"db6b0ed5-d5d3-463d-ae85-518a07d3c2b4"
- WAVOutputFormat
- MP3OutputFormat
- RAWOutputFormat
Prefer locale when you can. language only accepts base ISO codes like en. A request may set language or locale, never both.
en, fr, de, es, pt, zh, ja, hi, it, ko, nl, pl, ru, sv, tr, tl, bg, ro, ar, cs, el, fi, hr, ms, sk, da, ta, uk, hu, no, vi, bn, th, he, ka, id, te, gu, kn, ml, mr, pa, or, ur Prefer locale over language (for example en-GB). language only accepts base codes like en. locale also accepts regional codes like en-GB. Locale codes need Sonic 3.6+. A request may set language or locale, never both.
Text normalization. auto (default) runs the locale-aware normalizer, off skips it, or pass a locale code (for example en-IN) to pin the normalizer independently of the generation language. See Text Normalization.
The ID of a pronunciation dictionary to use for the generation. Pronunciation dictionaries are supported by sonic-3 models and newer.
Configure the various attributes of the generated speech. Available on sonic-3 and sonic-3.5; not available on earlier models.
See Volume, Speed, and Emotion for a guide on this option.
Response
Audio bytes
The response is of type file.