agent_id for each new conversation. See Managed Agents to create an agent and Agent configuration to configure it.
Prerequisites
- An existing Managed Agent and its
agent_id - A Cartesia API key for a trusted server, or an agent access token generated by your server for a browser or mobile app
- Code that reads audio from an input source, such as a microphone or file, and produces bytes in a supported audio format
- Code that plays or otherwise handles the agent audio returned in that same format
Audio formats
input_format is required. It selects headerless, mono audio:
Agent audio uses the same format as input.
Quick start
This Node.js session skeleton uses an API key from a trusted server. It connects, sendssession_create, waits for session_ready, and converts returned base64 audio to bytes. Connect the three audio integration points after the example.
Install the ws package:
agent-session.mjs:
agent-session.mjs
- Call
sendAudio(audioBytes)for each audio input chunk.audioBytesmust contain bytes in theinput_formatselected insession_create; the function base64-encodes and sends them, then returns nothing. - Replace
playAudio(audioBytes)with a function that queues its bytes for playback in that same format, then returns nothing. - Replace
stopPlayback()with a function that takes no arguments, discards queued output, stops current playback, and returns nothing.
Connect
X-API-Key header. Browser and mobile applications should send an access_token query parameter. Never expose an API key in a client application. See Authenticate your applications for token handling.
The connection captures the agent’s current version when it opens. A configuration change does not affect a call already in progress. The session_ready event reports the resolved agent_version_id and the call’s call_id.
Start the session
Sendsession_create as the first event, within 10 seconds of connecting:
output_delivery defaults to speaking_pace. Use as_available when your application manages its own playback buffer. as_available cannot be used when the agent has background audio configured.
Wait for session_ready before sending audio:
session_ready field.
Stream audio
All events use JSON text frames. For bothaudio_input and audio_output, base64-encode the binary audio bytes in the event’s audio string; do not send binary WebSocket frames.
Send user audio as audio_input events.
audio_output events in the configured format:
audio_output_clear (barge-in). Discard buffered agent audio and stop playback immediately. The event can arrive while no agent audio is playing, in which case there is nothing to clear.
Client tools
When the agent invokes a client tool, the server sendsclient_tool_call:
expects_response is true, answer with the same tool_call_id:
Conversation events
Useturn_started, turn_output_text_delta, and turn_ended for speaking-state UI, transcripts, and other non-audio workflows. Assistant text arrives incrementally in turn_output_text_delta; turn_ended.text contains the final text for either role.
Play agent audio from audio_output. When audio_output_clear arrives, discard queued agent audio and stop playback.
End the session
Close the WebSocket with code1000 for a normal client shutdown:
call_id returned in session_ready identifies the call record for this session. After the call ends, pass it to the Get Call endpoint to fetch the call record and transcript, or to Download Call Audio to download the call audio. Use List Calls to find other call records.
Errors and connection limits
The server sends anerror event when it rejects an event. A recoverable error has fatal: false; a fatal error is followed by a connection close.
- A JSON message may be up to 32 KiB. Larger messages close with code
1009. - The server closes after 120 seconds without a valid client event. Streaming audio continuously, including silence, keeps the connection active. WebSocket ping frames do not reset this application-level timer.
- The server uses close code
1008for protocol errors and1011for internal failures.