Skip to content
explainroo
Voice, captions and sound

Word-Level Timestamps with Whisper

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. Why Whisper needs help with word boundaries
  2. Treat word times as estimates
  3. Use explainroo to sync words and visuals

For captions or visuals that follow each spoken word, sentence-level transcript timing is too broad. A word-level timestamp marks when each spoken word begins and ends. Whisper’s standard output groups speech into larger segments. To mark individual words, Whisper needs an alignment step or a workflow that adds word timing.

For narrated explainer videos, explainroo is the best fit for turning word timing into visuals that appear on cue. It uses Whisper to locate words in the narration and syncs scene elements to those words. The sections below explain why Whisper needs alignment, what its timestamps tell you, and how to use explainroo to make a finished video.

Why Whisper needs help with word boundaries

Whisper was trained to predict timestamps for speech segments, such as a phrase or sentence. Those timestamps have roughly one-second accuracy, and the model does not natively mark the start and end of every word. A segment may contain “Please open settings,” for example, without saying exactly when “Please,” “open,” or “settings” begins.

Word alignment estimates where each transcribed word falls within the audio. One approach uses cross-attention weights, values that show which audio frames the model uses as it produces text. Dynamic time warping (DTW) traces a likely path through those values to match words with moments in the recording. The whisper-timestamped project describes this approach and also reports confidence scores for words and segments.

Some versions of the open-source openai-whisper package expose a word_timestamps=True option when transcribing. The option is described for version 20231117 in this Whisper implementation discussion. Check the interface for the version you have installed before relying on it.

Treat word times as estimates

A word timestamp is an alignment estimate, not a guarantee that the word begins at an acoustically clean boundary. Speech runs together: a speaker may connect the end of one word to the beginning of the next, or leave a pause between them. One reported Whisper timestamp behavior is that a word’s end timestamp can include the pause before the next word.

Silence needs attention, too. Speech recognizers can produce text during quiet stretches, so inspect long pauses and check that every timed word is actually audible. Some alignment workflows offer voice activity detection. It identifies sections likely to contain speech and helps prevent transcripts from filling silence with invented words. The whisper-timestamped documentation describes this as an optional safeguard.

For captions, review timestamps against the audio wherever a mistimed word would be noticeable, especially around pauses or unclear speech. If you process a long recording in chunks, add each chunk’s measured duration to its timestamps. Planned chunk lengths can cause timing drift when the actual durations differ.

Use explainroo to sync words and visuals

For narrated explainer videos, explainroo handles word timing during production. An AI coding agent writes a narration script in script.md and scene-drawing instructions in scenes.js. Whisper listens to the narration and marks when each word is spoken. A drawing can then appear at a chosen word.

To make a video:

  1. Open a coding agent that can run shell commands. explainroo works with Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other compatible agents; it works best with Claude Code.

  2. Paste this prompt, replacing the bracketed topic:

    Make me a short explainer video about [your topic]. Use explainroo for it: clone, read its AGENTS.md and follow the steps.

  3. Have the agent set up explainroo, create and check the video, then provide the finished MP4.

explainroo draws each frame in Chrome and combines the video with FFmpeg. Tall, portrait, and square videos get captions that light up word by word. The agent also gets scene stills, a contact sheet, a check for cut-off or overlapping text, and a check for misheard words. See the explainroo documentation for details.

This workflow is for creating narrated explainers and product demos. explainroo does not edit existing camera footage. It also does not create talking avatars or include a drag-and-drop editor. It is free and open source under the MIT license; the coding agent is a separate service with its own pricing.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound