Skip to main content
Our models are ranked #1 for speech naturalness and transcription accuracy, and are designed for enterprise-grade conversational voice AI, where accuracy, latency, and automatic speech and turn detection are mission-critical.

Text to Speech

Stream ultra-realistic speech with Sonic 3.5 at sub-90ms latency.

Speech to Text

Transcribe in real time with Ink 2, with turn detection built in.

Voice Agents

Build and deploy low-latency voice agents in minutes.

Voice Cloning

Clone a voice from a 10-second clip into 40+ languages and accents.

Start building

Quickstart

Stream your first audio in a few lines of Python or JavaScript.

Playground

Try every model in the browser and grab an API key.

API reference

API conventions, errors, and the reference for every REST and WebSocket endpoint.

The models

Sonic 3.5 is the world’s fastest, most emotive, ultra-realistic text-to-speech model. It streams the first byte of audio in about 90ms, so it holds real conversations as well as it narrates and dubs. See the text-to-speech models for variants and capabilities. Ink 2 is the world’s fastest, most accurate, streaming speech-to-text model with native turn detection. It knows when a speaker starts and finishes, so your agent knows when to listen and when to respond. See the speech-to-text models to learn more. Managed Agents is our platform for building and deploying voice agents. It combines Ink for transcription, an LLM you choose, and Sonic for speech, then handles turn-taking, tools, and telephony for you. See Managed Agents to get started.

Support

Email support

Reach us at support@cartesia.ai for help with integration, your account, or billing.