Docs
How it works
What happens on your computer between your request and the finished MP4.
What does a video project contain?
Every video is a folder with a few files:
videos/<name>/
video.json the look, size, voice, music and captions
script.md what the voice says, one "## scene" per scene
scenes.js one drawing function per scene
assets/ your images: screenshots, logos, illustrations
build/ the voice, the timing and the sound that explainroo makes
out/ video.mp4, stills and contact sheets
The agent writes three of these files. script.md holds the words the voice says. scenes.js is a bit of JavaScript that draws the pictures. video.json holds the settings. explainroo makes everything in build/ and out/.
Where does the voice come from?
Kokoro, an open voice model, reads the script aloud. It has 28 English voices. It runs on your computer’s processor through ONNX Runtime, the software that runs the model, so it needs no graphics card. You don’t need an account or an API key for it.
The voice speaks each sentence on its own. explainroo cuts the silence around each one and puts a short pause between sentences. That keeps the pace even.
How does a picture appear on the right word?
When Kokoro has spoken a scene, Whisper listens to the recording. It writes down when each word starts and ends. explainroo matches these words with the script, and every word in the script gets a time.
In scenes.js a picture can then appear on a word. With at: 'database' it shows up the moment the voice says “database”. This step also shows when the voice said a word differently from the script. That is how explainroo catches misread names and abbreviations.
Who draws the frames?
explainroo starts Chrome in the background and loads its own drawing code. For every frame, it calls the scene function with the time of that frame. The function then draws the whole picture. Each frame depends only on the time, so the same project always gives the same video.
The frames can hold hand-drawn lines, icons, charts and code. The lines come from Rough.js, and the 1,854 icons come from Lucide. The fonts come with explainroo.
Where do the music and the sound effects come from?
explainroo writes the music itself, as long as each video. It picks chords and instruments in a style that fits the look. The music gets quieter while the voice speaks. The sound effects follow what is on screen. A box that appears pops or makes a marker scribble, depending on the look. An arrow whooshes, and each list item plays a note.
Both the music and the sound effects are made in JavaScript. There are no sound files and no sample licenses.
How is the MP4 made?
explainroo draws the frames in several Chrome pages at once and hands them to ffmpeg. ffmpeg packs them into H.264 video. The sound is mixed on its own and set to a loudness of -14 LUFS. That is a common level for online video. Then ffmpeg joins the picture and the sound into one MP4.
How does the agent check the video?
An agent can’t watch a video. So explainroo shows it still pictures of every scene and contact sheets. It also checks the layout, for text that is cut off or on top of other text, and the pronunciation, for words the voice got wrong. All of that happens before the final video. The checking page explains each check.
How long does it take?
On a laptop with an 8-core Intel i9, the 68-second DNS example takes about 50 seconds to turn into a 1080p video once its voice exists. Making the voice takes a little less time than the narration lasts. When you change the script, only the changed scenes get a new voice.
What leaves my computer?
explainroo’s own part runs on your computer. The voice, the timing and the drawing send nothing anywhere and cost nothing. The speech models are downloaded once from Hugging Face. Set EXPLAINROO_OFFLINE=1 if you want to be sure no model files are fetched after that.
Your coding agent is a separate service with its own terms and costs. AI images are optional. If you use them, your image prompts go to OpenRouter.