Skip to main content
You can create a custom voice from a short recording in seconds using Cartesia’s instant voice clone. 10 seconds of audio is enough to get started, and up to 60 seconds to better retain the speaker’s accent.
The 60-second limit applies to Sonic 3.6 and newer. Older models only learn from the first 10 seconds of your clip, so a longer sample won’t improve results there.
Uploaded audio files must be no larger than 16 MB.

Record your clip

The clone picks up the mood of your clip. A cheerful take gives you a voice that sounds cheerful in everything it says, so record in the mood you want the voice to have.
  • One speaker, no background noise. A quiet room, no music, no echo off bare walls.
  • Speak the way you want the clone to speak. Tone, accent, pacing, and energy all carry over.
  • Speak naturally, not like an announcer. Over-enunciation and theatrical delivery make the clone sound less conversational.
  • Record in a language the speaker actually speaks. Each clone is built from one language. To make the clone speak other languages natively, clone it in its own language first, then add languages with the Add Voice Accents API.

Create the clone

You can create an instant voice clone in the Cartesia dashboard or through the API.
  • In the dashboard: open the instant voice clone tab, then record your clip directly or upload it as a file.
  • With the API: call Clone Voice with your clip, a name, and the language you recorded in.
An instant voice clone works well for most voices. If the speaker has a rare accent, is a character voice, or carries distinctive qualities you need to preserve, use a Pro Voice Clone instead. It trains on 30 minutes or more of audio from the same speaker, so it holds on to those details.