Skip to main content
Stream text to Cartesia over a WebSocket and receive audio chunks back in real time. This is ideal when text arrives incrementally, such as from an LLM, and it’s the pattern most voice agents use.

Prerequisites

  • A Cartesia API key. Create one here, then add it to your .bashrc or .zshrc:
    When accessing the Cartesia API from a browser, please use ephemeral access tokens for authentication to keep your API Key safe. See Authenticate Your Client Applications.
  • ffplay (part of FFmpeg), used to play audio output: Download the FFmpeg executable package for your operating system from the FFmpeg download page.
  • A language runtime and package manager:

Stream text and play audio

1

Install the client library

See the Cartesia Python client library for more details.
2

Stream text over a WebSocket

realtime-tts.py
3

Run the quickstart

This will stream text inputs to Cartesia, and play the streaming audio output using ffplay. (Make sure your device volume is turned on!)

How it works

The WebSocket connection can manage multiple contexts where each context is a full-duplex, continuous stream. You push text chunks in and receive generated audio chunks out in real time. This works well when generating text from an LLM in real time: Cartesia’s text-to-speech system maintains context history and appends each new chunk to it. This keeps generated speech continuous and consistent in tone and prosody while minimizing latency since you don’t have to wait for the full transcript to be ready. To summarize, here’s what our code does after establishing a WebSocket connection:
  1. Create a context with context().
  2. Push text incrementally with push(). Each call sends the chunk with continue: true, telling the model more text will follow. See continuations for details.
  3. Signal completion with no_more_inputs(), which sends continue: false to tell the model no more text is coming.
  4. Receive audio chunks as they are generated.
This uses a similar streaming pattern to realtime LLM WebSocket APIs: send text fragments as they arrive, and receive generated output incrementally — audio chunks in this case.

What’s next

Pick a voice

Choose a voice, or clone your own, then copy the voice ID back into this quickstart.

Tune the request

Change voice, model_id, or output_format, then rerun and compare output quality and behavior.

Stream continuations

Send incremental text while preserving flow across chunks for smoother long-form or LLM-driven speech.