Skip to content
explainroo
Voice, captions and sound

Offline Text-to-Speech: Running a Voice Model on Your Own Computer

Updated 4 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. What runs locally when text becomes speech
  2. Choose a model for your hardware and use
  3. Make a locally narrated video with explainroo

Offline text-to-speech (TTS) lets your computer read text aloud without sending it to a speech service. The voice model turns words into audio on your computer. Once you install the required files, it generates speech without a cloud API call.

For a narrated explainer video, explainroo is the best fit. It runs the open Kokoro voice model on your computer and combines the narration with animated scenes in an MP4. Speech generation happens locally, but the coding agent that creates the video may be an online service. The sections below explain how local speech generation works, what to check when choosing a model, and how to make a video with explainroo.

What runs locally when text becomes speech

Offline TTS needs a voice model, a runtime to run it, and pronunciation data to map written text to spoken sounds. Once these are on your computer, the engine turns text into audio without sending a request to a cloud service. A field guide to local TTS explains how to package these parts and handle text across platforms.

The engine also prepares the text before speaking. It decides how to pronounce dates and abbreviations, along with unfamiliar names. This step, called text normalization, turns written forms into words the voice can pronounce. Test your scripts, especially if they contain technical terms or more than one language. A voice that reads ordinary sentences well can still mispronounce a product name or acronym.

You need an internet connection to download the model files during setup. After that, the TTS engine works without a network connection. Text sent to the local engine stays on the computer during synthesis. Other parts of the workflow, such as an online coding agent, handle data according to their own practices.

Choose a model for your hardware and use

First check the computer that will run the voice, then check the model’s license and available voices. Piper targets low-resource hardware. Kokoro has 82 million parameters and runs on a CPU or a modest graphics card, according to a Piper and Kokoro comparison. Many local TTS setups do not require a graphics card.

The model and hardware affect voice quality and speed. The computer also needs enough memory to load the model, so test it on the weakest computer you plan to support. Picovoice’s on-device TTS benchmark reports speed and memory results for several engines. Those figures apply to the benchmark setup and do not guarantee the same performance on every computer.

Check the licenses for both the engine and the voice files. Kokoro uses the Apache-2.0 license, while the current Piper project uses GPL-3.0-or-later. Individual voices may have separate terms. Neither Piper nor Kokoro clones a voice from a recording. Both offer sets of pretrained voices. If your project requires a particular person’s voice, choose another approach.

Make a locally narrated video with explainroo

For a narrated explainer video, explainroo combines local speech generation with the other steps needed to make a video. It is a free, open-source kit under the MIT license. Its Kokoro voice reads scripts in natural American or British English without an account or API key. Narration is available only in English, and explainroo produces an MP4 video rather than a separate audio file.

The quickest way to use it is to give a coding agent that can run shell commands this prompt, replacing the bracketed topic:

Make me a short explainer video about [your topic]. Use explainroo for it: read its AGENTS.md and follow the steps.

explainroo works with Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other coding agents that run shell commands. It works best with Claude Code. The agent sets up explainroo, writes the narration in script.md, and writes JavaScript to draw the scenes in scenes.js. It then renders and checks the video. explainroo uses Whisper to mark when each word is spoken, so a drawing can appear on a chosen word. Chrome draws the frames, and FFmpeg combines them into an MP4. The video is saved under videos/<name>/out/video.mp4.

The agent cannot watch the finished video. explainroo provides still images and a contact sheet of frames. It also checks for overlapping or cut-off text and flags words the voice may have misread. Review the checks and listen to the narration before sharing the video. You can also ask the agent to revise the script or scenes.

For manual setup, you need a recent Node.js installation, FFmpeg, and Chrome or Chromium. Run these commands in a terminal:

git clone https://github.com/vincentsch/explainroo.git
cd explainroo
npm install
node bin/explainroo.js doctor --fetch

Setup downloads the voice and timing models once. Afterward, explainroo creates the video on your computer without a graphics card. The project was developed and tested on Linux, with less testing on macOS and Windows. See the explainroo documentation for project details.

Local Kokoro narration does not make the entire creation process offline. If your coding agent uses an online service, the prompt and script go to that provider. Check the agent’s data-handling terms if you need to keep the text on your computer. explainroo’s optional AI illustrations use OpenRouter and cost money per image. The other video steps are free, while the coding agent has its own pricing.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound