Skip to main content
Pronunciation dictionaries let you specify how to pronounce specific words and phrases. A dictionary is a simple search and replace, which directs the model to use another string in lieu of the text from the transcript. The pronunciation can be either an IPA pronunciation or a “sounds-like” guidance:
Save these JSONs as pronunciation dictionaries through our API or through our playground: Once a dictionary is created, use it in any TTS API by passing its id as pronunciation_dict_id. With the dictionary above, the string I ate some jambalaya on tchoupitoulas street becomes I ate some <<ˈ|dʒ|ə|m|ˈ|b|ə|ˈ|l|aɪ|ˈ|ə>> on chop-uh-TOO-liss street before being handed off to the model.

Case sensitivity

Each entry has a case_sensitive flag. The default is false. When case_sensitive is false, new york matches New York, NEW YORK, and any other capitalization. Case-insensitive matching requires Sonic 3.6+. When case_sensitive is true, a lowercase key also matches sentence-start capitalization (cat / Cat, not CAT). Mixed-case keys match only that spelling (LaTeX does not match latex). The API rejects two entries that would collide under these rules, for example latex and LaTeX when either is case-insensitive. Use false for common words when you want every capitalization. Use true when case distinguishes words, such as LaTeX (typesetting) vs latex (rubber).

Sharing a dictionary

New dictionaries are private by default, so only your organization can use them. To use this pronunciation dictionary for Text-to-Speech generation in an external account, set its access to public. To see which dictionaries are currently shared, call the List Pronunciation Dictionaries API and look for entries with access.type set to public.