Text-to-Speech (SSE)
Stream audio with extra metadata from a complete transcript
Authorizations
Cartesia API key (sk_car_...). Get one at play.cartesia.ai/keys.
Headers
API version header.
2026-08-14 "2026-08-14"
Body
The ID of the voice.
"db6b0ed5-d5d3-463d-ae85-518a07d3c2b4"
The transcript's language or locale (for example en or en-GB). language and locale accept the same values. Set one or the other, never both; setting both returns a 400 error. See supported codes.
The transcript's language or locale (for example en or en-GB). locale and language accept the same values. Set one or the other, never both; setting both returns a 400 error. See supported codes.
Usually unnecessary: Cartesia picks the closest accent the voice supports for the requested language or locale. Set it only to make a multilingual voice sound accented (e.g. speak English with a French accent). Must come from the voice's Get Voice accents field. Learn more here.
Text normalization. auto (default) runs the locale-aware normalizer, off skips it, or pass a language or locale code (for example en or en-IN) to pin the normalizer independently of the generation language. See Text Normalization.
Whether to return word-level timestamps. If false (default), no word timestamps will be produced at all. If true, the server will return timestamp events containing word-level timing information.
Whether to return phoneme-level timestamps. If false (default), no phoneme timestamps will be produced. If true, the server will return timestamp events containing phoneme-level timing information.
Whether to use normalized timestamps (True) or original timestamps (False).
The ID of a pronunciation dictionary to use for the generation. Pronunciation dictionaries are supported by sonic-3 models and newer.
Configure the various attributes of the generated speech. Available on sonic-3 and newer models; not available on earlier models.
See Volume, Speed, and Emotion for a guide on this option.
This can be any string value you find useful. The server will echo back the same context_id in events that it sends.
Contexts on the TTS (WebSocket) endpoint are used for continuations. > The TTS (SSE) endpoint does not support continuations, so most users just ignore this property.
Response
Server-sent events stream. Each frame is data: <json>\n\n where the JSON payload matches TTSSSEEvent.
- TTSSSEChunkEvent
- TTSSSETimestampsEvent
- TTSSSEPhonemeTimestampsEvent
- TTSSSEDoneEvent
- TTSSSEErrorEvent
Audio data chunk.
Event type identifier.
chunk Whether this is the final event for the request. Always false for chunk events.
false Base64-encoded audio data.
Server-side processing time for this chunk in milliseconds.
HTTP-style status code.
The context ID echoed back from the request, if one was provided.