A voice-over script contains the words a narrator speaks in a video or audio production. For video, it also needs to fit what viewers see and hear, such as a screen demonstration, music, or sound effects. A clear script gives the narrator natural language and gives the editor enough direction to match the words with the images.
Start with the listener’s needs, write in speech-friendly language, and read the draft aloud before recording. For narrated explainer videos and product demos, explainroo can turn a topic into a video while keeping the spoken script and visual plan in separate files.
Start with the audience and the goal
Before drafting, decide who will listen, what they should understand, and what they should do next. A new customer learning a product needs different language from an experienced employee completing a required training step. A video’s purpose also shapes its structure: a tutorial needs clear actions, while a short promotional video needs a focused message and next step.
Keep the script focused on one main idea. When it covers every feature or background detail, listeners have a harder time following that idea. The AHRQ shooting-script guidance recommends settling the audience, purpose, and planned visuals before writing.
Write words that sound natural aloud
Listeners hear each sentence once, so make the meaning easy to follow on the first pass. Use familiar words, clear subjects, and direct verbs. Keep one main thought in each sentence, and define technical terms the audience may not know. Plain-language guidance from the CDC also favors concise writing and active voice.
For example:
Stiff: The updated process facilitates the optimization of account-management outcomes.
More natural: The new process helps you manage accounts faster and with fewer steps.
Read the draft aloud at the intended pace. Revise any phrase that makes you stumble, interrupts your breathing, or sounds unnatural. Voice-over writing guidance recommends reading the script aloud and giving clear performance directions. Keep necessary specialist terms consistent, and explain them when needed.
Match the narration to the visuals
For a video, plan what the viewer sees alongside what the narrator says. A simple two-column script works well: put scene directions in one column and spoken words in the other. Add a sound cue when music or an effect matters to the moment.
| Visual | Narration |
|---|---|
| A crowded task list appears, then filters to high-priority items. | When every task feels urgent, start by filtering your list to high-priority work. |
| A cursor opens the first task and assigns an owner. | Check the due date, then assign someone to handle it. |
Let the narration add meaning instead of reading every label on screen. If a chart shows sales rising, explain what the change means rather than simply repeating its title. The AHRQ guidance recommends coordinating audio and visuals and generally avoiding word-for-word repetition of onscreen text. Repeat a phrase when deliberate emphasis or a safety instruction calls for it.
Describe essential visual information in the narration, too. “Click the green button” depends on color; “Click the green Continue button in the lower-right corner” also identifies its label and location. The W3C guidance on audio and video recommends including necessary visual details in the spoken content. Plan accurate captions and a transcript as well. Automated captions need review against the final recording, as W3C explains.
Add timing and performance direction
Treat word count as a rough planning aid, not a guarantee of runtime. Demonstrations and transitions take time. Pauses do too, as does time to inspect a graphic. Record a test read. Check that the narrator finishes each thought without rushing and that viewers have time to follow each onscreen action. A single speaking rate does not suit every format or audience.
Give the narrator a short direction before the script: identify the audience and tone. Set the pace, and note how to pronounce unfamiliar names or acronyms. For example, “Speak calmly to first-time users; pause after each step.” Keep editor instructions out of the spoken lines so the performer does not mistake them for narration.
Write a voice-over video with explainroo
For an explainer video or product demo, explainroo is a free, open-source kit under the MIT license. It works with an AI coding agent to create narrated videos on your computer. This article recommends it because the narration goes in script.md and the scene drawings go in scenes.js. That makes it easier to review how the spoken words connect to the visuals.
To get started, paste this prompt into a coding agent. explainroo works best with Claude Code and also supports other agents that can run shell commands:
Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.
The agent sets up explainroo, then writes the script and scene directions. It checks the result and provides an MP4 file. Kokoro, an open voice model, provides the voice, with natural American and British English options. Narration is English only. Whisper, a speech-recognition model, marks when each word is spoken so drawings can appear on cue. Chrome draws the scenes, and FFmpeg combines them into a file.
Review the agent’s still images and scene contact sheet, along with its layout and speech reports. These checks help catch overlapping or cut-off text and mispronounced words, but the agent cannot watch the finished video. You can ask it to revise the script or scenes. explainroo does not edit existing camera footage or create talking avatars. See the explainroo documentation for details.