Skip to main content
This guide covers migrating from ElevenLabs Realtime Speech to Text when used with commit_strategy=vad.

All migration guides

This guide contains both bare API descriptions and SDK code. To install the SDK:
If you’re already using the Cartesia SDK, upgrade to version >=3.2.0
Ink 2 only supports English right now.
We expect to add more languages in the coming months.

Connection

Replace the ElevenLabs WebSocket URL and auth header with Cartesia’s /stt/turns/websocket.
In browsers, WebSockets do not support request headers. Instead, pass the API version as the cartesia_version query param and use a short-lived access token using the access_token query param instead of an API key. Connect to the auto-finalization WebSocket with the Cartesia SDK:

Query parameters

ElevenLabs Scribe (VAD)Cartesia Realtime STT (Auto)Notes
model_id=scribe_v2_realtime requiredmodel=ink-2 requiredSee Models for all options.
audio_format=pcm_16000encoding=pcm_s16le + sample_rate=16000 requiredElevenLabs bundles format and rate; Cartesia splits them. See encoding.
commit_strategy=vadSee manual finalization for manual commits.
language_codeink-2 only supports en right now. More languages are coming soon!
cartesia_version=2026-03-01 requiredSee API Conventions for details.
vad_silence_threshold_secs, vad_threshold, min_speech_duration_ms, min_silence_duration_msCartesia uses semantic turn detection. No VAD tuning required.
include_timestampsComing soon!
keytermsComing soon!
enable_loggingControlled by your organization.
ElevenLabs bundles the sample format and rate into a single audio_format token. Cartesia splits them into encoding and sample_rate.
ElevenLabs audio_formatCartesia encodingCartesia sample_rate
pcm_8000pcm_s16le8000
pcm_16000pcm_s16le16000
pcm_22050pcm_s16le22050
pcm_24000pcm_s16le24000
pcm_44100pcm_s16le44100
pcm_48000pcm_s16le48000
ulaw_8000pcm_mulaw8000
Cartesia also accepts pcm_s32le, pcm_f16le, pcm_f32le, and pcm_alaw.
All Cartesia encodings support all sample rates.

Sending audio

ElevenLabs wraps each audio chunk in a JSON formatted text frame and base64-encodes the audio bytes.
Cartesia accepts audio chunks as binary frames: send the raw audio bytes directly:
  • No need to supply previous text
  • Sample rate is determined upon connection by the sample_rate query parameter
If you currently commit audio mid-session with ElevenLabs, consider using Cartesia with manual finalization instead.Take a look at the migration guides page for details.
To commit all audio and close the session, send a JSON formatted text frame:
Cartesia will transcribe all buffered audio, then close the socket for you.

Sending audio with the SDK

Decoding base64 encoded audio before sending

Committing and closing

Event mapping

Scribe emits a partial_transcript, then a committed_transcript when its VAD commits a segment. Cartesia folds the same information into a turn lifecycle: turn.start, turn.update, turn.eager_end, turn.resume, and turn.end. See Turn Detection for the full state machine.
ElevenLabs message_typeCartesia typeNotes
session_startedconnectedConnection confirmed. You do not need to wait for it before sending audio.
partial_transcriptturn.updatePartial transcript while the user is speaking.
committed_transcriptturn.endUser stopped speaking; contains the complete transcript for the user turn.
committed_transcript_with_timestampsturn.endTimestamps are not yet available.
turn.startThe user began speaking. Carries no transcript.
turn.eager_endThe model predicts the user might be done speaking. Okay to ignore.
turn.resumeThe user kept talking; ignore the last turn.eager_end.
errorerrorClient or server errors.
auth_errorCartesia will reject the WebSocket upgrade with a 401 or 403 HTTP status.
quota_exceedederrorCartesia’s error response will contain "error_code": "quota_exceeded".
rate_limitederrorCartesia’s error response will contain "error_code": "concurrency_limited".
session_time_limit_exceededCartesia will send a WebSocket close frame with code 1001.

Partial transcripts

An ElevenLabs partial_transcript:
Becomes a Cartesia turn.update:

Committed transcripts

An ElevenLabs committed_transcript:
Becomes a Cartesia turn.end:

Example Server Messages

Scribe’s transcripts are joined with spaces. Ink’s are not.
ElevenLabs Scribe (VAD)Cartesia Realtime STT (Auto)
turn.start
partial_transcript "Scribe's transcripts"turn.update "Scribe's transcripts"
turn.eager_end "Scribe's transcripts"
turn.resume
partial_transcript "Scribe's transcripts are joined with spaces."turn.update "Scribe's transcripts are joined with spaces."
turn.eager_end "Scribe's transcripts are joined with spaces."
committed_transcript "Scribe's transcripts are joined with spaces."turn.end "Scribe's transcripts are joined with spaces."
committed_transcript_with_timestamps "Scribe's transcripts are joined with spaces."
turn.start
partial_transcript "Ink's are not."turn.update " Ink's are not."
turn.eager_end " Ink's are not."
committed_transcript "Ink's are not."turn.end " Ink's are not."
committed_transcript_with_timestamps "Ink's are not."

References

API Reference

Cartesia Realtime STT (Auto)

Full Code Example

Using the Cartesia SDK