If you need spoken narration for a project, Kokoro is an open-weight text-to-speech (TTS) model, software that turns written text into audio. Its 82 million parameters, the numerical values it uses to generate speech, make it relatively compact. Run it through Python or use it as the voice engine in a larger project.
For narrated explainer videos, explainroo uses Kokoro for English narration and handles the visuals, timing, music, and finished MP4. For standalone speech generation, start with Kokoro’s official Python package.
What Kokoro does
Kokoro generates speech from text in a selected language and voice; its model weights, the files containing its trained parameters, are released under the Apache license. The official repository describes the 82-million-parameter model as fast and cost-efficient compared with larger models. These comparisons come from the project, so test your text and setup before using them to make a production decision.
Kokoro v1.0, released on January 27, 2025, supports eight languages and 54 voices, according to its Hugging Face model card. The language codes include a for American English and b for British English. It also supports Spanish and French. The other supported languages are Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese. The chosen voice affects how the narration sounds; for example, the repository’s Python example uses the American English voice af_heart.
The model card says Kokoro uses a few hundred hours of audio and phoneme labels. Phonemes are the speech sounds that make up words. Its documented training sources include public-domain audio. They also include audio licensed under Apache or MIT terms, as well as synthetic speech from commercial TTS models.
Run Kokoro from Python
The official repository shows how to install the Python package and use it for inference. Inference means using a trained model to generate an output, here audio, from text.
Install the package and audio-file library:
pip install "kokoro>=0.9.4" soundfile
Kokoro also relies on espeak-ng for English fallback pronunciation and for some non-English languages. Install it separately when your environment requires it. Windows users need a separate installer.
This example creates audio with American English settings and writes each generated segment as a WAV file:
from kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code="a")
text = "Kokoro turns written text into spoken audio."
for index, (graphemes, phonemes, audio) in enumerate(
pipeline(text, voice="af_heart")
):
sf.write(f"segment-{index}.wav", audio, 24000)
The output sample rate is 24,000 samples per second. The pipeline yields audio segments, so the example saves each segment separately. For longer narration, check how the package divides text and join the segments if you need a single audio file.
Choose between local use and a hosted service
Local inference runs the model on your machine, so you control the text and generated audio. You also manage the Python environment, model files, and dependencies. Hosted services handle some setup, but their providers set the pricing, privacy, and availability terms. Open Kokoro model weights do not make every service that offers Kokoro free or affiliated with the model’s authors.
Use the official repository and model page as your starting points. The model page warns that sites using “Kokoro” in their root domain are not affiliated with the model or its author. Treat claims from unofficial sites, including claims about extra languages or compatible APIs, as unverified unless the official project confirms them.
Fast TTS generation alone does not guarantee a responsive voice service. In a live call, speech generation adds delay alongside speech recognition, network requests, and other components. Production services also need safeguards for concurrent requests and fallback behavior. A voice-pipeline analysis discusses these layers. Test the complete service under realistic conditions instead of judging it with a standalone generation test.
Make a narrated video with explainroo
For an explainer video or product demo, explainroo uses Kokoro to read the script and produces a finished MP4. It is a free, open-source kit under the MIT license. You direct a coding agent, which writes the narration in script.md and the scene instructions in scenes.js; explainroo renders the video on your computer.
To get started, paste this prompt into a coding agent that can run shell commands, replacing the bracketed topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
The agent sets up explainroo, creates and checks the video, then gives you an MP4 file. explainroo works with Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other coding agents that can run shell commands. It works best with Claude Code. The agent is a separate service and may have its own costs.
Kokoro provides the English narration, with a choice of natural American and British English voices and no separate account or API key for the voice. Whisper, a speech-recognition model, listens to the recording and marks when each word is spoken, so scene elements can appear in sync with the narration. Chrome draws the video frames on an HTML canvas. Scenes can include charts, code, or screenshots, and explainroo adds background music and sound effects.
The agent cannot watch the rendered video. explainroo provides still frames, a contact sheet, a check for cut-off or overlapping text, and a check for mispronounced words. In video.json, choose from paper, clean, chalk, blueprint, and midnight looks. Set the format to wide, tall, portrait, or square. Tall, portrait, and square videos include word-by-word captions.
For product demos, tell the agent which product to show and where to find its code or website. The agent rebuilds the product screens with matching colors, fonts, and button labels, then scripts a pointer to click and type while the narration explains the product. explainroo does not edit existing camera footage, create talking avatars or live-action footage, or provide a drag-and-drop editor. Ask the coding agent to make changes.
For manual installation, explainroo requires recent Node.js, FFmpeg, and Chrome or Chromium. Its setup command downloads the voice and timing models once; no graphics card is needed. The project is developed and tested on Linux, with less testing on macOS and Windows. The installation instructions cover setup, and the example videos show the output.