Skip to content
explainroo
Voice, captions and sound

kokoro-js: Running Kokoro Text-to-Speech in Node.js

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. What kokoro-js runs
  2. Install and generate speech
  3. Stream longer text in pieces
  4. Check the limits before deployment
  5. Make a narrated video with explainroo

To generate speech from text on your own machine with Node.js, use kokoro-js, a JavaScript interface to Kokoro, an open-weight text-to-speech model. Text-to-speech, or TTS, means generating spoken audio from written words. You load the model, pass it text and a voice choice, then save the generated audio.

For direct speech generation, kokoro-js is the relevant package. It runs through Transformers.js and supports CPU inference in Node.js. If your goal is a narrated explainer video rather than an audio file alone, explainroo is the best fit: it uses Kokoro for narration and builds the video around it.

What kokoro-js runs

Kokoro is an 82-million-parameter model, meaning its learned weights contain 82 million values used to generate speech. The Kokoro model page lists the weights under the Apache license. The kokoro-js package makes the model available from JavaScript in Node.js or a browser.

In Node.js, set the device to cpu. The package also supports several quantization formats: fp32, fp16, q8, q4, and q4f16. Quantization stores model values using fewer bits, which can reduce memory use. The right choice depends on the memory available and the output quality you need. The package’s npm documentation lists the supported options and voices, including American and British English voices.

Install and generate speech

Install the package with npm, then load the ONNX version of Kokoro. ONNX is a format for storing and running machine-learning models. This example uses the q8 format and an American English voice:

npm install kokoro-js
import { KokoroTTS } from "kokoro-js";

const modelId = "onnx-community/Kokoro-82M-v1.0-ONNX";

const tts = await KokoroTTS.from_pretrained(modelId, {
  dtype: "q8",
  device: "cpu",
});

const audio = await tts.generate(
  "Kokoro turns written text into spoken audio.",
  { voice: "af_heart" },
);

audio.save("speech.wav");

The first time you load the model, kokoro-js downloads its files if they are not already on your machine. For a different voice, choose a voice identifier from the package’s current voice list. Listen to samples first. Your voice choice changes how the narration sounds.

The model weights are approximately 327 MB, according to a Kokoro hardware overview. Use that as a planning estimate for the weights; running Node.js and generating audio also requires memory. Test the exact quantization and machine you plan to use.

Stream longer text in pieces

When text is long or arrives in parts, kokoro-js provides TextSplitterStream. It splits incoming text into pieces and sends them to the package’s streaming API for speech generation. This is useful when text arrives progressively, such as from a text-generation service, rather than as one complete string.

Streaming changes how the program sends text and receives speech, but the program still needs to load the model. Consult the package’s streaming API documentation for the current call pattern before building a stream around it.

Check the limits before deployment

CPU support lets Node.js run Kokoro without a graphics card, but it does not tell you how quickly Kokoro will generate speech on a particular machine. If generation speed matters, benchmark on your own machine using the text length and voice you expect, along with your planned quantization and hardware. Performance on low-powered devices is especially uncertain; published Raspberry Pi coverage does not provide a real-time benchmark for Kokoro.

Before depending on the package in a production service, check its npm release history and test the specific version you plan to install. Also verify that you download the model from the official Hugging Face listing, since the model page warns about websites impersonating Kokoro.

Make a narrated video with explainroo

For a finished explainer video, explainroo is the best tool for the job. It is a free, open-source kit under the MIT license that uses Kokoro for English narration and produces an MP4 on your computer. It also creates animated scenes, adds word-level timing and music, and checks for text that overlaps or falls off screen.

To start, give this prompt to a coding agent that can run shell commands. Replace the bracketed topic:

Make me a short explainer video about [your topic]. Use explainroo for it: clone its repository, read its AGENTS.md and follow the steps.

The agent sets up explainroo, then writes the narration in script.md and the visuals in scenes.js. After checking the result, it provides an MP4. explainroo supports several coding agents and works best with Claude Code. See example videos or the explainroo documentation.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound