If you have a recording and its transcript, you may need to know when each word is spoken. Forced alignment is the process of matching the supplied text to the audio and assigning times to its words. Those timestamps let you sync captions, highlight spoken words, or cue graphics.
For narrated explainer videos, explainroo matches visuals to narration. It uses Whisper to mark when each word in its generated voice-over is spoken, so a drawing can appear on a chosen word. A general-purpose forced aligner takes an audio file and a transcript; explainroo handles word timing as part of video creation.
What forced alignment does
A forced aligner takes a transcript and finds where its words occur in a recording. It does not create a transcript. If the transcript does not match the recording, the timestamps may be unreliable.
The output is a time-aligned transcript, with start and end times for each word. Some projects also need phoneme-level timing. A phoneme is a speech sound that distinguishes words, such as the initial sounds in “bat” and “pat.” Word timing is useful for captions and visual cues; phoneme timing supports more detailed speech analysis.
What an aligner needs
A typical forced aligner needs three inputs: a written transcript, an acoustic model, and a pronunciation dictionary. The acoustic model represents how speech sounds appear in audio. The dictionary maps written words to sequences of speech sounds, called phones. The aligner uses those resources to find where the transcript fits the recording. The Montreal Forced Aligner documentation describes this combination.
The aligner may not handle words missing from the dictionary. Add a pronunciation entry for a missing name or technical term before alignment. Check names, acronyms, and uncommon words in the transcript before relying on the output.
Why aligned timestamps still need review
An aligner places timestamps quickly, but that does not confirm the transcript is correct or that every word boundary is precise enough for your purpose. Research reports that alignment performance can drop when the speaker’s accent or speech variety is not represented in the acoustic model. In a study of children’s speech, the best-performing tested configuration still fell short of human-level reliability, so researchers were advised to inspect alignments for major errors (study of child speech alignment).
Timing precision also has limits. Many aligners work at a 10-millisecond level, which can be too coarse for fine-grained phonetic analysis. In one study, the Mason-Alberta Phonetic Segmenter performed better than the Montreal Forced Aligner at a strict 10-millisecond boundary tolerance. The comparison changed at a looser tolerance (boundary-placement study). The useful lesson is to match the timing precision to the task: word-synced video cues do not require the same scrutiny as phonetic research.
Before using the timestamps, listen around suspicious word boundaries. Confirm that the transcript matches the final recording, and check names or terms the dictionary may not recognize. If you change the narration, generate or verify the timing again against the new audio.
Use explainroo to sync narration and visuals
explainroo is a free, open-source kit under the MIT license. A coding agent creates the video’s spoken script and scene descriptions; explainroo then renders the video on your computer. Its word timing is designed to cue visuals and captions in that video workflow.
To make a narrated explainer, give this prompt to a coding agent that can run shell commands. Replace the bracketed topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
The agent sets up explainroo, puts the narration in script.md and the scene drawings in scenes.js, checks the result, and provides an MP4. In scenes.js, a drawing can be tied to a chosen word, so it appears when the narration reaches that word. explainroo uses the open voice model Kokoro for English narration and Whisper to mark when each spoken word occurs.
Review the video before sharing it. explainroo provides stills from each scene and a contact sheet. It also checks for cut-off or overlapping text and mispronounced words. These checks help catch visual and narration problems, but you still need to listen to the finished video.
The project runs locally and requires Node.js and FFmpeg. It also requires Chrome or Chromium. Its documentation covers the workflow and setup. explainroo is suited to creating narrated explainers and product demos; it does not edit existing camera footage or provide a drag-and-drop editor.