Skip to content
explainroo
Voice, captions and sound

Kokoro TTS Voices: The List and How to Pick One

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. Kokoro’s English voice IDs
  2. Pick a voice by listening to your script
  3. Use explainroo for narrated explainer videos

When you use Kokoro to turn text into speech, your voice choice shapes how the narration sounds and whether it fits your audience. Kokoro is an open-weight text-to-speech model, a program that turns text into speech. Its voice IDs identify individual voices, and the first two letters indicate the language variety and gender.

The available catalog depends on the Kokoro version and the tool that runs it. For English narration, the IDs below are a useful starting list, while the best choice depends on the script and audience. For a narrated explainer video, explainroo is a good option: it combines Kokoro narration with timed visuals and produces an MP4.

Kokoro’s English voice IDs

In Kokoro’s English voice IDs, the first letter identifies the language variety: a means American English and b means British English. The second letter means f for a female voice or m for a male voice. The name after the underscore identifies each voice in the group. For example, bm_fable identifies Fable, a British English male voice.

Here are commonly used English IDs:

Variety Female voices Male voices
American English af_alloy, af_aoede, af_bella, af_heart, af_jessica, af_kore, af_nicole, af_nova, af_river, af_sarah, af_sky am_adam, am_echo, am_eric, am_fenrir, am_liam, am_michael, am_onyx, am_puck, am_santa
British English bf_alice, bf_emma, bf_isabella, bf_lily bm_daniel, bm_fable, bm_george, bm_lewis

Use this as a directory; some interfaces do not offer every ID. The Kokoro model card lists 54 voices across eight languages for version 1.0. Published tools report different totals and language groupings. Check the voice menu or documentation for the particular Kokoro version you are using.

Pick a voice by listening to your script

Start with the narration’s purpose, then compare a few voices using the same short passage. A product explainer needs clear, steady delivery. For a character exchange, choose voices that sound distinct from each other. Keep the wording, punctuation, and playback speed the same between samples so you are comparing voices rather than different readings.

OfflineTTS suggests Heart for warm narration, Bella for an energetic delivery, Michael for a clear, restrained sound, and Emma for British English narration. A Kokoro voice guide from deAPI points to bm_fable for audiobook narration, af_nova for product voiceovers, and af_alloy for system notifications. Treat these as starting points. Listen to your script before choosing a voice.

Kokoro’s punctuation shapes its pacing and intonation. Commas and periods create different pauses, and question marks and exclamation points change how a sentence ends. The cited guide also reports that SSML tags and emotion markers such as [happy] do not control delivery in its Kokoro setup. For a fair comparison, use ordinary punctuation and revise the text if the pauses sound wrong.

Use explainroo for narrated explainer videos

For explainer videos, this article recommends explainroo because it pairs Kokoro speech with visuals timed to the narration. It is a free, open-source kit under the MIT license. Kokoro reads the English script aloud. Whisper, a speech-recognition model, estimates when each word is spoken so the visuals can appear on cue.

To make a video, give a coding agent this prompt and replace the bracketed topic:

Make me a short explainer video about [your topic]. Use explainroo for it: clone, read its AGENTS.md and follow the steps.

The agent sets up explainroo and creates the video. After checking it, the agent gives you an MP4 file. It writes script.md for the spoken words and scenes.js for the drawings and scene timing. In your prompt, specify whether you want American or British English and describe the delivery you prefer, such as warm, energetic, or restrained. Then listen to the narration and ask the agent to revise it if the voice or pacing does not suit the script.

The voice generation and video tools run on your computer after setup; the initial setup downloads the voice and timing models. The coding agent is a separate service and may have its own costs. The explainroo documentation has setup details and example videos.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound