{
"type": "session_create",
"audio": {
"input_format": "pcm_44100",
"output_delivery": "speaking_pace"
}
}{
"type": "audio_input",
"audio": "base64_encoded_audio_data"
}{
"type": "dtmf_input",
"digit": "5"
}{
"type": "client_tool_result",
"tool_call_id": "call_2b7e4f9a1c0d",
"result": "2 items in cart",
"is_error": false
}{
"type": "session_ready",
"call_id": "ac_gqkgRWUz2u64qFUjA1mZyr",
"agent_id": "agent_Fo7pKNBUwLZxrTd6jvhpaE",
"agent_version_id": "av_7Hq2mXbK9cLdNfPzR3tWvE",
"audio": {
"input_format": "pcm_44100",
"output_delivery": "speaking_pace"
}
}{
"type": "audio_output",
"audio": "base64_encoded_audio_data"
}{
"type": "audio_output_clear"
}{
"type": "dtmf_output",
"digit": "5"
}{
"type": "client_tool_call",
"tool_call_id": "call_2b7e4f9a1c0d",
"tool_name": "open_cart",
"parameters": {
"cart_id": "cart_456"
},
"expects_response": true
}{
"type": "turn_started",
"turn": 3,
"role": "user",
"start_time": 12.48
}{
"type": "turn_output_text_delta",
"turn": 4,
"role": "assistant",
"text": " world"
}{
"type": "turn_ended",
"turn": 3,
"role": "user",
"text": "I'd like to check my order status.",
"interrupted": false,
"start_time": 12.48,
"end_time": 15.02,
"tool_calls": []
}{
"type": "error",
"code": "invalid_event",
"message": "dtmf_input must be one of 0-9, * or",
"fatal": false
}WebSocket
Stream audio and events between an application and a Managed Agent. Server applications authenticate with X-API-Key; browser and mobile applications pass an access_token query parameter. See the WebSocket guide for setup, limits, and audio handling.
{
"type": "session_create",
"audio": {
"input_format": "pcm_44100",
"output_delivery": "speaking_pace"
}
}{
"type": "audio_input",
"audio": "base64_encoded_audio_data"
}{
"type": "dtmf_input",
"digit": "5"
}{
"type": "client_tool_result",
"tool_call_id": "call_2b7e4f9a1c0d",
"result": "2 items in cart",
"is_error": false
}{
"type": "session_ready",
"call_id": "ac_gqkgRWUz2u64qFUjA1mZyr",
"agent_id": "agent_Fo7pKNBUwLZxrTd6jvhpaE",
"agent_version_id": "av_7Hq2mXbK9cLdNfPzR3tWvE",
"audio": {
"input_format": "pcm_44100",
"output_delivery": "speaking_pace"
}
}{
"type": "audio_output",
"audio": "base64_encoded_audio_data"
}{
"type": "audio_output_clear"
}{
"type": "dtmf_output",
"digit": "5"
}{
"type": "client_tool_call",
"tool_call_id": "call_2b7e4f9a1c0d",
"tool_name": "open_cart",
"parameters": {
"cart_id": "cart_456"
},
"expects_response": true
}{
"type": "turn_started",
"turn": 3,
"role": "user",
"start_time": 12.48
}{
"type": "turn_output_text_delta",
"turn": 4,
"role": "assistant",
"text": " world"
}{
"type": "turn_ended",
"turn": 3,
"role": "user",
"text": "I'd like to check my order status.",
"interrupted": false,
"start_time": 12.48,
"end_time": 15.02,
"tool_calls": []
}{
"type": "error",
"code": "invalid_event",
"message": "dtmf_input must be one of 0-9, * or",
"fatal": false
}ID of the managed agent to connect to.
API version, e.g. 2026-08-14
Use an API key when you're calling from a trusted server.
Use a short-lived access token when calling from a browser or client app. Learn more here.
Configures the session's audio. This must be the first message sent. The server closes the connection if any other event arrives first, or if no event arrives within 10 seconds. Unknown fields are rejected.
Streams user audio to the agent. Send chunks continuously (for example every 20–100 ms) for the best latency. audio must be base64-encoded audio in the format declared in session_create.audio.input_format. Wait for session_ready before sending audio; audio sent earlier is dropped.
Sends a DTMF (dual-tone multi-frequency) digit to the agent.
Answers a client_tool_call whose expects_response is true. tool_call_id must match the call being answered. The decoded result may be at most 4 KiB; larger results are replaced with a result_too_large error result. Late, duplicate, or mismatched results are ignored.
Sent once after session_create, when the agent pipeline can accept audio. Reports the call record created for this session and the agent version the session is pinned to. Use call_id with the calls API to fetch the recording and transcript after the call.
The agent's speech. audio is base64-encoded audio in the format declared in session_create.audio.input_format.
Sent when the user starts speaking. Discard buffered agent audio and stop playback. The event can arrive while no agent audio is playing, in which case there is nothing to clear.
A DTMF digit the agent sends, for clients bridging a telephony system.
The agent asks this client to run a client tool. When expects_response is true, answer with a client_tool_result carrying the same tool_call_id; the invocation stays open until a result, error, or timeout. When false, dispatching the action completes the invocation.
Sent when the user starts speaking or the agent starts responding. turn comes from a single counter shared by both roles, starting at 1 and strictly increasing over the call, so turn alone identifies a turn.
Sent as the agent speaks, carrying the text spoken since the previous delta, typically one word at a time. Deltas include separator spaces, so build the running text by appending text verbatim. The user's speech is not streamed incrementally. It arrives as finalized text in turn_ended.
Sent when a turn finishes, with the complete final text for the turn. This is the version to store and display. The events for a single turn always arrive in order: turn_started, then any turn_output_text_delta events, then turn_ended. If the agent hangs up mid-turn, the connection closes without a final turn_ended; treat the close as ending any open turn.
Reports a problem with the session or an event you sent. When fatal is false, the offending event was dropped and the stream stays open. When fatal is true, the server closes the connection: code 1008 for client and protocol errors, 1011 for agent pipeline failures.
Was this page helpful?