In a video with word-by-word animated captions, each spoken word lights up or changes style as the voice reaches it. To create the effect, match each scripted word to its exact place in the audio, then update the captions as the video plays.
For narrated explainer videos, explainroo creates the narration, visuals, and word-by-word captions for tall, portrait, and square videos. Whether you use explainroo or make captions yourself, the text and word timing both need to be accurate.
How the words stay in sync
The software that displays captions needs each spoken word and its start and end times. It uses the times to highlight each word, move to the next, or show a phrase.
One way to get the timings is forced alignment. This process takes a known transcript and matches its words to the audio. Automatic speech recognition (ASR) converts speech into text and helps find the alignment. NVIDIA’s explanation of forced alignment describes how the method matches the reference text to the speech instead of testing every possible transcript.
For example, suppose the narration says, “Captions follow the spoken words.” The timing data might mark “Captions” from one point in the audio to the next, then do the same for “follow,” “the,” “spoken,” and “words.” When playback reaches each interval, the video highlights the matching word.
The transcript must match what the speaker actually says. If it contains a typo or leaves out a spoken word, the timing can drift or attach to the wrong word. Check the audio against the transcript before using the animation.
How captions appear during playback
Caption timing can be stored in a text track, which holds captions with start and end times, or built into the video image. WebVTT is a plain-text format for timed captions. A player can use those timestamps to display captions in sync with the audio.
For word-by-word captions, the display also needs to know which word is being spoken. Caption software can show a phrase and change the current word’s style when its time comes. Another approach is to render the caption animation directly into the video frames, so it appears in the image wherever the video plays. A separate caption track is still useful for viewers who need selectable captions.
Animated highlighting does not make captions accurate or easy to read. Check that the words are correct, readable, and synchronized with speech. Captions also support access and understanding. A review of more than 100 empirical studies found benefits for attention, memory, and understanding, including for people who are deaf or hard of hearing and people learning a language (review of caption research). Accessibility guidance also emphasizes accurate, synchronized captions (caption guidance).
Make word-by-word captions with explainroo
explainroo is a free, open-source kit under the MIT license. It works with coding agents that can run shell commands and is designed to create narrated explainer videos and product demos. To make a captioned explainer, ask the agent to use explainroo and name the subject and video format.
Paste this prompt into a supported coding agent, replacing the bracketed text:
Make me a short explainer video about [your topic]. Use explainroo for it: clone explainroo, read its AGENTS.md and follow the steps.
The agent sets up explainroo, makes and checks the video, then provides an MP4. It writes a script.md file for the narration and a scenes.js file that describes the visuals. You can assign a drawing to a word in the script so the visuals change along with the narration.
Whisper listens to the recorded narration and marks the start and end of each spoken word. explainroo uses those timings to cue visuals and produce word-by-word captions in tall, portrait, and square formats. It uses Kokoro, an open voice model with American and British English options. Narration is available in English only.
Check the generated video before accepting it. explainroo provides still images of scenes, a contact sheet of frames, a layout check for overlapping or cut-off text, and a speech check for mispronounced words. Correct any transcript or narration errors, then ask the agent to revise the video. The kit cannot edit existing camera footage and does not include a drag-and-drop editor. Instead, describe the video to the agent and request changes.