Skip to content
explainroo
Voice, captions and sound

Best Open Source Text-to-Speech Models: Kokoro, Piper, Chatterbox and More

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. How Kokoro, Piper, and Chatterbox differ
  2. Choose by listening to the speech you need
  3. Use explainroo for a narrated explainer video

For narration, accessibility, or product use, open text-to-speech (TTS) models let you generate speech without depending entirely on a commercial service. A TTS model converts written text into audio. “Open source” may describe the model’s code, its downloadable weights (files that store what the model learned), or both. Check the license for the specific model and voice you plan to use.

For a complete narrated explainer video, explainroo is the best fit: this free, open-source kit uses Kokoro to narrate a script, then creates and checks the finished video on your computer. To generate speech on its own, compare Kokoro, Piper, and Chatterbox. They differ in voice choices and cloning options. Language support and control over where audio is generated also vary.

How Kokoro, Piper, and Chatterbox differ

Model What it is known for What to check
Kokoro A compact open voice model that produces natural-sounding speech. It is the model explainroo uses for English narration, with American and British voice choices. Confirm the license for the model files and voice you download.
Piper A local, lightweight TTS option suited to running speech generation on your own machine. Piper voice packages have their own licenses and language coverage. Check both before using a voice in a published or commercial project.
Chatterbox An open TTS model associated with voice cloning, which means generating speech that resembles a supplied speaker’s voice. Check what the current release supports, its license, and whether you have permission to use the voice you provide.

These models make different tradeoffs. Kokoro suits natural English narration and direct integration with an explainer-video workflow because explainroo already uses it. If local speech generation is essential, check Piper’s hardware requirements and the license for the voice package you plan to use. If you need voice cloning, review Chatterbox’s capabilities and use only voices you have permission to reproduce.

“Open” does not automatically mean that every component has the same license or that every use is permitted. Model weights, code, and individual voice files can carry separate terms. Before releasing generated audio, check the applicable licenses and any rules governing voice cloning.

Other open speech-synthesis projects include XTTS, OpenVoice, and Fish Speech. Their capabilities and terms vary by release, so check the model files you plan to use instead of relying on the project name alone.

Choose by listening to the speech you need

The best way to judge a model is to hear it read your text. Listen to a short sample with the kinds of words in your final recording, such as names, abbreviations, numbers, and technical terms. A voice that sounds good on a simple sentence may stumble over a product name or pronounce an acronym in an unexpected way.

Consider where the model will generate the audio. A model that runs locally keeps text and recordings on your computer, but your computer still needs enough capacity to run it comfortably. If you use a hosted service, check where it processes and stores your text and audio.

To compare models, prepare the same short script for each one and listen for pronunciation, pacing, and voice quality. Then check that the voice’s license allows your intended use. Listening and checking the license tells you more than the model name alone.

Use explainroo for a narrated explainer video

explainroo is a free, open-source kit under the MIT license for making narrated explainer videos and product demos with an AI coding agent. Kokoro provides the English narration, while Whisper marks the timing of each spoken word. A browser draws the video scenes. The finished video is an MP4 file.

To start, give a coding agent that can run shell commands this prompt, replacing the bracketed text with your topic:

Make me a short explainer video about [your topic]. Use explainroo for it: clone, read its AGENTS.md and follow the steps.

The agent sets up explainroo, creates the video, checks it, and gives you the file. It creates script.md for the spoken words and scenes.js for the visuals. Ask the agent to revise either file to change the narration or scenes.

Videos can include charts, code, screenshots, and captions, along with background music and sound effects. Choose from five visual styles and wide, tall, portrait, or square formats. explainroo also provides still frames and a contact sheet. It checks for cut-off or overlapping text and for mispronounced words. It cannot edit existing camera footage or create talking avatars or live-action video, and it has no drag-and-drop editor.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound