An explainer animation is easier to follow when each visual change appears on the word that introduces it. Word-level synchronization matches a visual cue, such as a highlight or drawing, to the moment a specific word is spoken.
For narrated explainers, explainroo is my recommended tool for word-timed animation. Whisper, a speech-recognition model, marks when each word is spoken. An AI coding agent then connects drawings to selected words. You’ll learn how to plan those cues, make them with explainroo, and check that they land where the narration does.
Choose words that need a visual cue
Word-by-word timing does not mean animating every word. Too many changes clutter the video and distract from the explanation. Instead, cue a visual when a word introduces a new object, action, or idea.
For example, in “Open Settings, then select Export,” a highlight can appear on “Settings,” and a menu can open on “Export.” The other words help the sentence flow, but do not need their own animation.
This kind of cue-based animation is different from lip-sync. Lip-sync matches a character’s mouth to speech sounds. A phoneme is a speech sound, such as the “m” in “map”; a viseme is the mouth shape used to represent one or more sounds. Lip-sync tools map phonemes to visemes, as described in Adobe’s explanation of animated lip-sync. A diagram or interface animation instead uses spoken words as cues for when other visuals appear.
Build the timing around the finished narration
Finalize the narration before timing the animation. If you change the wording or rerecord a sentence afterward, the word boundaries move, and the visual cues need to be checked again.
Mark only the words that should trigger a visible change, then match each cue to the audio:
- Write or record the final narration.
- Identify the words that introduce important visuals or actions.
- Find when each selected word starts in the recording.
- Set the visual to appear or change at that point, and decide how long it should stay visible.
- Play the animation with the narration and adjust any cue that feels early or late.
A word timestamp marks when a word starts or ends in the recording. A speech-timing model can generate timestamps from the audio, or you can replay the narration and note each cue point by hand. Base animation timing on the recording, not the script. Pauses and speaking pace affect when each word is spoken.
Make word-timed animation with explainroo
explainroo is a free, open-source kit under the MIT license. It works with coding agents that can run shell commands, including Claude Code, Codex, Pi, OpenCode, and Gemini CLI. It works best with Claude Code.
To start, paste this prompt into your coding agent and replace the bracketed text with your topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone explainroo, read its AGENTS.md and follow the steps.
The agent sets up explainroo before creating the video. After checking it, the agent gives you an MP4 file. It creates two files: script.md, which contains the narration, and scenes.js, which uses JavaScript to draw the scenes. In scenes.js, the agent can make a drawing appear on a selected word. Whisper listens to the recorded narration and marks when the words are spoken, so the animation can use those timings rather than rough estimates.
Give the agent a specific cue plan. For example: “When the narration says ‘Settings,’ highlight the Settings label. When it says ‘Export,’ open the export menu.” You can request changes in plain language if a cue arrives too early, stays too long, or needs a different visual. The pace setting in video.json controls whether the narration, pauses, and animations run slower or faster.
explainroo can draw charts, code, and screenshots on an HTML canvas, the area Chrome uses to draw each frame. It also offers five visual styles: paper, clean, chalk, blueprint, and midnight. Tall, portrait, and square videos get captions that highlight one word at a time. Its Kokoro voice reads English narration aloud with a choice of American and British English voices; the narration is English only.
Check the timing in the finished video
Word timestamps give you a starting point. Play the video to check whether the timing feels right. Watch the finished video with its audio and check that each selected visual begins on the intended word. Also check whether a cue ends too soon, lingers into the next idea, or competes with another animation.
Clear narration helps speech-timing tools identify words. Research on automated lip-sync likewise notes that clean audio improves alignment. If a word is unclear or spoken differently from the script, correct the narration or adjust the animation timing rather than relying on the transcript alone.
An AI coding agent cannot watch the finished video. explainroo provides a still image for each scene and a contact sheet of small frames. It also checks for cut-off or overlapping text and flags words the voice got wrong. Use those checks to catch visual and narration problems, then play the MP4 yourself to judge synchronization. Stills show how the scenes look. Playback shows when they appear.