Skip to main content
Our models are ranked #1 for speech naturalness and transcription accuracy, and are designed for enterprise-grade conversational voice AI, where accuracy, latency, and automatic speech and turn detection are mission-critical.

Text to Speech

Stream ultra-realistic speech with Sonic 3.5 at sub-90ms latency.

Speech to Text

Transcribe in real time with Ink 2, with turn detection built in.

Voice Agents

Build and deploy low-latency voice agents with Line.

Voice Cloning

Clone a voice from a 10-second clip into 40+ languages and accents.

Start building

Quickstart

Stream your first audio in a few lines of Python or JavaScript.

Playground

Try every model in the browser and grab an API key.

API reference

API conventions, errors, and the reference for every REST and WebSocket endpoint.

The models

Sonic 3.5 is the world’s fastest, most emotive, ultra-realistic text-to-speech model. It streams the first byte of audio in about 90ms, so it holds real conversations as well as it narrates and dubs. See the text-to-speech models for variants and capabilities. Ink 2 is the world’s fastest, most accurate, streaming speech-to-text model with native turn detection. It knows when a speaker starts and finishes, so your agent knows when to listen and when to respond. See the speech-to-text models to learn more. Line is our platform for building and deploying voice agents. It brings voice to your text agents and handles audio orchestration, deployment, and observability with Sonic and Ink built in, so you can focus on your agent’s reasoning. See Voice Agents to get started.

Support

Email support

Reach us at support@cartesia.ai for help with integration, your account, or billing.