A voice-over is narration recorded separately from the picture and placed over video. You can record your own voice, generate narration, or record directly inside a video editor. In every case, match the spoken words to the visuals, keep the voice clear, and check the finished video for timing and sound problems.
If you are making an explainer video or product demo from scratch, explainroo uses an AI coding agent to create the narration and visuals and produce an MP4. To add your own voice to existing camera footage, use a video editor. explainroo does not edit existing footage.
Plan the narration around the visuals
Start with a rough sequence of video clips, even if you have not finalized every cut. A timeline is the editing workspace where video and audio clips are arranged over time. A playhead marks the moment currently selected. Watching the visuals first helps you write lines that match each scene.
Write a direct script that is easy to say aloud. Add useful information instead of repeating details viewers can already see. For example, “Click the blue button to open the export menu” is clearer than narrating the button’s location and appearance.
You can draft a simple two-column script:
| What viewers see | What the voice says |
|---|---|
| The timeline appears | “Move the playhead to where the narration should begin.” |
| Music starts under the video | “Lower the music while you speak.” |
Read the script out loud before recording. Revise any sentence that feels awkward to say or runs longer than the visual moment allows.
Record clean audio
Use a built-in phone or computer microphone, or connect a separate one. Select it as the recording input. Wear headphones or mute your speakers so the microphone does not pick up playback.
Record a short test using one of your loudest lines. Watch the input meter and listen back. Clipping occurs when the recording level is too high for the device to capture, making the sound harsh or distorted. Lower the input level and try again if the meter signals clipping or the recording sounds distorted. Audacity’s recording guidance suggests keeping peaks around -6 dB as headroom, but that is a practical recommendation for Audacity, not a universal delivery requirement.
You can record directly on an audio track in your video editor, or record separately and import the audio file. An audio track is a timeline lane that holds sound clips. Keeping narration on its own track makes it easier to trim and adjust. Apple’s iMovie instructions show how to record narration at the playhead and add it to the timeline as an editable clip.
Place the voice and balance the sound
Put the narration on its own track and line up each phrase with the matching visual. Trim false starts and long pauses. Split the clip when the subject changes or the narration needs to match a new scene.
Make the voice easy to understand before adding music. Ducking lowers background music while someone speaks. Adjust the music by ear, or use an editor’s ducking feature if available. Noise-reduction tools can help with background sound, but heavy processing can make a voice sound thin or unnatural. Compare the cleaned recording with the original and reduce the effect if the voice sounds worse.
Add captions after you finish editing the narration. Captions include speech and relevant non-speech sounds, while subtitles may include dialogue alone. The W3C caption guidance recommends keeping captions synchronized and positioned so they do not cover important visual information.
Make a narrated video with explainroo
For an explainer or product demo made from scratch, explainroo automates the narration, scene visuals, timing, and video assembly. It is a free, open-source kit under the MIT license. An AI coding agent runs it on your computer; explainroo works with agents that can run shell commands and works best with Claude Code.
Give the agent this prompt, replacing the bracketed words with your subject:
Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.
The agent sets up explainroo, creates the video, checks it, and gives you an MP4. It writes script.md, which contains the spoken words, and scenes.js, which describes the visuals. The open voice model Kokoro reads the script in a natural American or British English voice; narration is English only, with no account or API key needed. Whisper, a speech-recognition tool, marks when each word is spoken so scene changes can follow the narration.
The video runs locally on your computer. You can choose a visual style and video format, and tall, portrait, and square videos get word-by-word captions. The agent checks scene stills, a contact sheet, text layout, and speech for errors. Optional AI illustrations cost money per image; other explainroo costs are zero per video, while the coding agent is a separate service with its own pricing. See the explainroo docs for details.
Export and listen again
Watch the full video after exporting. Check that the voice matches the visuals, every line is clear over the music, and the captions match the final narration. Listen on headphones and a phone speaker. YouTube warns that stereo audio can sound wrong when mobile playback combines it into mono. If you use music, confirm its license covers your intended use.