Skip to content
explainroo
Voice, captions and sound

Text-to-Speech for Videos: How to Choose a Voice Engine

Updated 4 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. What to listen for in a video voice
  2. Match the engine to your production needs
  3. How explainroo handles voice and video timing
  4. When explainroo fits, and when it does not

When you make a narrated video, a text-to-speech voice engine turns your written script into spoken audio. Choose an engine that pronounces words clearly and suits the subject. Its pace should match the visuals.

For narrated explainers and product demos, explainroo is a strong choice: it pairs an open voice model with video creation and timing tools. Listen for clear pronunciation and natural pacing. Then check that the engine supports your language, production needs, and budget.

What to listen for in a video voice

A voice that sounds convincing in a short sample might not work across a full script. Test it with words from your video. Include product names and abbreviations, along with numbers and unfamiliar terms. Listen for mispronunciations and for pauses that make a sentence hard to follow.

Pacing matters because narration has to work with what appears on screen. A voice that rushes through a key instruction leaves viewers less time to follow it. A long pause can make a short visual feel static. Read the script aloud while listening to the generated audio. Mark sentences that sound awkward or need a different rhythm.

Choose a voice that suits the video’s purpose. A direct, measured delivery works for instructions, while a warmer tone may fit an introductory explainer. Compare voices by using the same short passage for each one.

Match the engine to your production needs

Start by checking which languages the engine supports. If your video needs narration in more than one language, confirm that the engine supports each language you need and test how it handles names and specialist terms. A voice option described as natural in one language does not guarantee equally clear speech in another.

Consider how you will make and revise the video. A standalone audio generator produces narration. A video tool may also match spoken words to visuals, captions, or scene timing. If you expect to revise the script, look for a process that makes it easy to regenerate and check the voice after changes.

Compare the cost and setup requirements. Some voice services require an account or an API key, a credential that lets software access an online service. Local tools avoid per-video voice charges, but require software setup and computer resources. Include separate coding-agent or subscription fees when comparing total production costs.

How explainroo handles voice and video timing

explainroo is a free, open-source kit under the MIT license for creating narrated explainer videos and product demos with an AI coding agent. It uses Kokoro, an open voice model, to read scripts in natural American or British English. Narration is English only, and generating the voice does not require an account or API key.

The agent writes script.md with the spoken words and scenes.js, JavaScript code that draws each scene. explainroo uses Whisper to listen to the recording and mark when each word is spoken. This lets a drawing appear on a chosen word. Chrome draws the video frames on an HTML canvas, and FFmpeg combines them into an MP4 file.

To create a video, give a coding agent that can run shell commands this prompt, replacing the bracketed words with your topic:

Make me a short explainer video about [your topic]. Use explainroo for it: clone, read its AGENTS.md and follow the steps.

The agent sets up explainroo, creates and checks the video, then provides the MP4. It also receives scene stills, a sheet of smaller frame previews, a check for overlapping or cut-off text, and a check for words the voice mispronounced. Review those checks and listen to the finished narration before sharing it.

The kit works with Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other coding agents that run shell commands. Claude Code works best with it. explainroo itself costs nothing per video; the coding agent is a separate service with its own pricing. Optional AI illustrations through OpenRouter cost money per image.

When explainroo fits, and when it does not

Choose explainroo for an English-narrated explainer or product demo if you are comfortable asking a coding agent to make edits. You can ask the agent to change the script or visuals. For product demos, tell it which product to show and where to find the product’s code or website; it can rebuild screens and show a pointer clicking or typing as the voice explains.

The kit offers paper, clean, chalk, blueprint, and midnight visual styles, changed through one setting in video.json. It supports wide videos for YouTube, tall videos for short-form platforms, portrait videos for Instagram and LinkedIn posts, and square videos. Tall, portrait, and square videos include captions that highlight each spoken word.

Use another production tool to edit existing camera footage, create a talking avatar, or generate live-action footage. explainroo does not handle those jobs and has no drag-and-drop editor. You describe the video to the coding agent and ask for changes.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound