A video transcript is a readable text version of a video’s content. It includes spoken words and, when needed, sounds and visual details that help explain the video. A transcript is read separately; captions are timed text that appears alongside the video during playback.
For an existing video, draft the transcript, check it against the recording, and describe visuals that add information. If you are making a narrated explainer from scratch, I recommend explainroo. It creates the spoken script and uses speech timing to cue the visuals. It does not transcribe existing camera footage.
Choose a transcript, captions, or both
Use a transcript when viewers need a readable, searchable version of the video. Use captions when they need to follow speech and meaningful sounds while watching. Captions include speaker changes and laughter. They also identify significant sound effects, as described in W3C guidance.
If charts, on-screen instructions, or actions add information that the narration leaves out, make the transcript descriptive. Add concise notes about those visuals. W3C recommends a descriptive transcript when people need both audio and visual information to understand the video. For many prerecorded videos, captions and a transcript meet different needs.
Make and check a first draft
Start with the clearest available copy of the video. Draft the transcript by listening and typing, or use automatic speech recognition (ASR), software that turns speech into text. ASR saves typing, but its output needs human review. W3C warns that automatic captions alone are not accurate enough to count as accessible captions.
Listen to the recording as you edit. Check names, technical terms, numbers, dates, and links. Be especially careful when speakers overlap or the audio is unclear. Add speaker labels when they help readers follow the conversation. Mark sounds that affect meaning, such as [applause] or [alarm beeps], and describe essential visuals, for example: [On screen: A chart shows sales rising from Q1 to Q2.]
For a tutorial, break the transcript into short sections at topic or scene changes. Add timestamps when they help readers find a moment in a longer recording. A transcript can be plain text; timestamps alone do not make it a caption file.
Turn the text into captions when needed
Captions must match the timing of the video. A common web format is WebVTT, a text file that groups words into timed sections. The player uses the timings to show each section during playback. Check the finished captions in the actual player to catch timing errors and text that covers important content.
Place the transcript near the video so viewers can find it. Read it without playing the recording to make sure it makes sense on its own. Then play the video with captions to check synchronization and speaker labels.
Use explainroo for a new narrated explainer
For a new explainer video, explainroo creates a script you can use as the narration transcript. It is a free, open-source kit that works with coding agents, including Claude Code, Codex, Pi, OpenCode, and Gemini CLI. The agent creates script.md with the spoken words and scenes.js to draw the visuals. Whisper listens to the narration and notes when each word is spoken, so visuals can appear at the right time.
To try it, replace the brackets with your subject and give this prompt to a coding agent:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
The agent sets up explainroo, creates and checks the video, and gives you an MP4. You can review script.md as the narration transcript. You need a coding agent that can run shell commands. The agent is a separate service with its own pricing. See explainroo’s documentation for details.