Skip to main content
Connect your application to an existing Managed Agent over one bidirectional WebSocket to send user audio, receive agent audio, and handle conversation or client-tool events. This page assumes you already created and configured an agent. Reuse its agent_id for each new conversation. See Managed Agents to create an agent and Agent configuration to configure it.

Prerequisites

  • An existing Managed Agent and its agent_id
  • A Cartesia API key for a trusted server, or an agent access token generated by your server for a browser or mobile app
  • Code that reads audio from an input source, such as a microphone or file, and produces bytes in a supported audio format
  • Code that plays or otherwise handles the agent audio returned in that same format

Audio formats

input_format is required. It selects headerless, mono audio: Agent audio uses the same format as input.

Quick start

This Node.js session skeleton uses an API key from a trusted server. It connects, sends session_create, waits for session_ready, and converts returned base64 audio to bytes. Connect the three audio integration points after the example. Install the ws package:
Save this as agent-session.mjs:
agent-session.mjs
From a browser or mobile app, never embed an API key. Mint a short-lived agent access token on your server and pass it as a query parameter — browsers cannot set WebSocket headers.
Run the skeleton:
To complete the audio path:
  • Call sendAudio(audioBytes) for each audio input chunk. audioBytes must contain bytes in the input_format selected in session_create; the function base64-encodes and sends them, then returns nothing.
  • Replace playAudio(audioBytes) with a function that queues its bytes for playback in that same format, then returns nothing.
  • Replace stopPlayback() with a function that takes no arguments, discards queued output, stops current playback, and returns nothing.

Connect

Server applications authenticate with the X-API-Key header. Browser and mobile applications should send an access_token query parameter. Never expose an API key in a client application. See Authenticate your applications for token handling. The connection captures the agent’s current version when it opens. A configuration change does not affect a call already in progress. The session_ready event reports the resolved agent_version_id and the call’s call_id.

Start the session

Send session_create as the first event, within 10 seconds of connecting:
output_delivery defaults to speaking_pace. Use as_available when your application manages its own playback buffer. as_available cannot be used when the agent has background audio configured. Wait for session_ready before sending audio:
See the Agent WebSocket API reference for every session_ready field.

Stream audio

All events use JSON text frames. For both audio_input and audio_output, base64-encode the binary audio bytes in the event’s audio string; do not send binary WebSocket frames. Send user audio as audio_input events.
The server returns audio_output events in the configured format:
When the user starts speaking, the server sends audio_output_clear (barge-in). Discard buffered agent audio and stop playback immediately. The event can arrive while no agent audio is playing, in which case there is nothing to clear.

Client tools

When the agent invokes a client tool, the server sends client_tool_call:
If expects_response is true, answer with the same tool_call_id:
Late, duplicate, and mismatched results are ignored. Results may be up to 4096 bytes.

Conversation events

Use turn_started, turn_output_text_delta, and turn_ended for speaking-state UI, transcripts, and other non-audio workflows. Assistant text arrives incrementally in turn_output_text_delta; turn_ended.text contains the final text for either role. Play agent audio from audio_output. When audio_output_clear arrives, discard queued agent audio and stop playback.

End the session

Close the WebSocket with code 1000 for a normal client shutdown:
The call_id returned in session_ready identifies the call record for this session. After the call ends, pass it to the Get Call endpoint to fetch the call record and transcript, or to Download Call Audio to download the call audio. Use List Calls to find other call records.

Errors and connection limits

The server sends an error event when it rejects an event. A recoverable error has fatal: false; a fatal error is followed by a connection close.
  • A JSON message may be up to 32 KiB. Larger messages close with code 1009.
  • The server closes after 120 seconds without a valid client event. Streaming audio continuously, including silence, keeps the connection active. WebSocket ping frames do not reset this application-level timer.
  • The server uses close code 1008 for protocol errors and 1011 for internal failures.

API reference

See the WebSocket API reference for every event and field.