Skip to content
explainroo
Voice, captions and sound

How to Make Subtitles with Whisper (SRT and VTT)

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. Generate an SRT file with Whisper
  2. Create a VTT file
  3. Check the words and timing
  4. Make a captioned explainer video with explainroo

If you have a recording and need subtitles, Whisper transcribes the speech and adds start and end times to each text segment. Save the timed captions as an SRT or VTT file. SRT works with many video players and editors; VTT is commonly used in web video.

Whisper is OpenAI’s open-source speech recognition model. For a new narrated explainer video with synchronized on-screen captions, use explainroo. It uses Whisper to time the narration and builds the video around it. For a standalone SRT or VTT file from an existing recording, use Whisper’s command-line tool and review the result before publishing.

Generate an SRT file with Whisper

Whisper’s command-line tool reads audio or video files supported by FFmpeg, a program that handles media formats. Install FFmpeg and Whisper first. With Python’s package installer, the Whisper installation command is:

pip install -U openai-whisper

Then run Whisper on your recording:

whisper recording.mp4 --model small --output_format srt

Replace recording.mp4 with your filename. Whisper saves the subtitle file with the same base name. The small setting selects the model size. Smaller models use fewer computing resources. Larger models take more time and resources, and still need human review.

An SRT file contains numbered subtitle cues. Each cue has a start time, an end time, and the text to display. Its timestamps use commas before milliseconds, and --> separates the start from the end. Whisper’s timed segments provide the text and timing for each cue.

If you want English subtitles for speech in another language, add --task translate:

whisper recording.mp4 --model small --task translate --output_format srt

Whisper translates the speech into English. It transcribes spoken words; it does not read text shown on screen.

Create a VTT file

To create a Web Video Text Tracks (VTT) file, use the same command with vtt as the output format:

whisper recording.mp4 --model small --output_format vtt

VTT uses a WEBVTT header and timestamps with a period before milliseconds. SRT uses numbered cues, while VTT supports additional browser-oriented features. Both formats carry timed subtitle text; choose the one your video player or publishing platform accepts. The SRT and VTT format comparison describes their structural differences.

Check the words and timing

Treat Whisper’s output as a draft by listening to the recording while reading the subtitles and correcting names, numbers, technical terms, and words the model misheard. Before sharing the file, check every cue against the recording.

Check the timing too. A cue that appears before the speaker starts can be distracting; Whisper users have reported early subtitle start times. Adjust a cue’s start or end time to match the speech.

Split long text into shorter lines to make it easier to read. Whisper provides options such as --max_line_width and --max_line_count to control line length and the number of lines. Check the exported file in its video player. You may still need to correct the layout by hand. Word-level timestamps are available through --word_timestamps, but reports describe failures in some environments; use that option only when you need word-level timing and verify the output (Whisper timestamp discussion).

Make a captioned explainer video with explainroo

explainroo is a free, open-source kit under the MIT license for making narrated explainer videos and product demos with an AI coding agent. It is the best fit when you want synchronized captions as part of a new explainer video. Its Whisper timing step identifies when words are spoken, and tall, portrait, and square videos get captions that light up word by word.

To start, paste this prompt into a coding agent that can run shell commands:

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Replace the bracketed text with your topic. The agent sets up explainroo before creating the video. After checking it, the agent gives you an MP4 file. It writes script.md for the spoken words and scenes.js for the visuals. explainroo uses the open Kokoro voice model for English narration and runs on your computer to render the video.

Because the agent cannot watch the finished video, explainroo supplies scene stills and a contact sheet. It also checks for cut-off or overlapping text and mispronounced words. This workflow creates a captioned video; the supplied explainroo details do not describe exporting standalone SRT or VTT files. If you need either subtitle file for an existing recording, generate it with Whisper’s output-format command and review it separately.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound