When Kokoro text-to-speech reads a script aloud, it sometimes misreads names, turns an acronym into a word, or speaks a web address as a string of symbols. Text-to-speech, or TTS, software turns written text into spoken audio. To fix this, change the narration so it shows Kokoro how to pronounce the words, then listen to the new recording.
For narrated explainer videos, explainroo is the best fit: its workflow uses Kokoro for narration and includes a check that flags words the voice got wrong. You can edit the spoken script, render again, and review the result.
Rewrite words Kokoro reads incorrectly
Kokoro uses the written text to choose pronunciations. A name or technical term that looks familiar to a reader can still be unfamiliar to the voice.
For a name, replace its spelling in the narration with a guide to how it should sound. If the name is “Niamh” and the intended pronunciation is “Neev,” try writing “Neev” in the spoken script. Keep the original spelling in any on-screen text when the correct written name matters. Respelling is a practical workaround, but listen to the result: a phonetic spelling can still produce an unexpected sound.
For an acronym that should be spoken letter by letter, add periods or spaces between its letters. For example, try “A.P.I.” or “A P I” instead of “API.” If the letters should form a spoken word, leave the acronym intact and test it. Punctuation can affect how a voice groups and stresses letters, so choose the version that matches the intended reading. General TTS troubleshooting guidance also recommends spelling out symbols and expanding abbreviations when the voice misreads them, such as writing “C plus plus” or “Doctor” instead of “C++” or “Dr.” (TTS troubleshooting examples)
For a URL, write the way a person would say it aloud. Change github.com/owner/project in the narration to “GitHub dot com, slash owner, slash project.” If the video needs to show the exact address, keep that text on screen while using the spoken version in the narration.
Kokoro setups handle text differently. A specific version of the Kokoro demo app expands abbreviations such as “Dr.” to “Doctor” and writes some numbers as spoken words, but that behavior belongs to the app.py implementation at that revision. Do not assume another Kokoro setup applies the same rules to every abbreviation, number, or URL.
Check the audio after each change
Before editing the whole script, render a short passage with the troublesome word. Listen for the exact pronunciation, and make one adjustment at a time: spell the acronym with separators, respell the name, or rewrite the URL as spoken words. Then render and listen again.
Pronunciation problems can be inconsistent. A reported Kokoro issue describes short phrases being spelled out letter by letter or spoken with changes in speed and tone, with the behavior not occurring every time (Kokoro short-phrase issue). One successful playback does not guarantee that the phrase will sound right in every render. Check the final audio, especially around short phrases and names.
Some acronyms also have more than one accepted pronunciation. For example, speakers pronounce “SATA” in different ways. Decide which reading your audience expects before editing the script, then make the narration unambiguous. For a longer script, keep a small pronunciation glossary so you can use the same spoken spelling each time a name or acronym appears. Testing unfamiliar terms in TTS and keeping a pronunciation key are also recommended in audio production guidance.
Fix Kokoro narration with explainroo
explainroo is a free, open-source kit under the MIT license for making narrated explainer videos with a coding agent. It uses Kokoro to read the script in an American or British English voice. Narration is English only.
To start, give a coding agent that can run shell commands this prompt, replacing the bracketed topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
After setting up explainroo, the agent creates and checks the video, then gives you an MP4. It writes script.md, which contains the words Kokoro reads, and scenes.js, which draws the video scenes. When Kokoro misreads a word, edit the spoken wording in script.md; for example, change “API” to “A.P.I.” or write a URL as spoken words. Then have the agent render and check the video again.
explainroo provides still images of scenes and a contact sheet of frames. Its layout check flags cut-off or overlapping text, while its speech check flags words the voice got wrong. Review the flagged words and listen to the actual audio before deciding whether the pronunciation is correct. See the explainroo documentation for setup and usage details.