A voice-over is recorded narration that plays over an explainer video’s visuals. It explains the subject, sets the editing pace, and gives the video a human tone. A strong voice-over helps viewers understand and trust the message. A weak one can make clear visuals feel rushed, confusing, or too promotional.
Choose a voice that fits the audience, script, visuals, and distribution plan. The narrator might be a professional voice actor, an in-house speaker, or a modern synthetic voice. Test the actual script in context. Then protect the recording with clear audio, accessibility support, and written usage rights.
Define what the voice needs to accomplish
Start with the viewer, not a voice actor’s demo reel. Decide what the audience should understand, feel, or do after watching. A technical onboarding video needs clear explanations and a steady pace. A consumer product video may work better with a warmer, more conversational delivery.
Tone means the social and emotional quality of the performance. “Calm and precise” gives a narrator different direction from “warm and energetic.” Avoid vague instructions such as “sound professional.” Instead, describe the desired relationship with the viewer: informed guide, reassuring expert, helpful peer, or calm authority.
Also consider where the video will appear. A product page, paid advertisement, internal training video, and social clip may need different energy and licensing terms. Record these decisions in a short voice brief that includes:
- audience and subject knowledge;
- desired tone and traits to avoid;
- target runtime and language needs;
- pronunciation of product names and technical terms;
- visual references, such as a storyboard or rough animation;
- platforms, territories, duration, and future versions.
Fix the script before casting
A narrator cannot fix an overlong or confusing script through performance alone. Read the copy aloud and remove sentences that sound unnatural when spoken. Explain unfamiliar terms, shorten long clauses, and add pauses after major ideas.
For a clear informational explainer, start at about 130 to 150 words per minute. Energetic product videos may use a faster pace, but speed should not make up for too much copy. U.S. Section 508 guidance warns that speech above roughly 180 words per minute can make captions hard to read and synchronize. Use these figures as production guidance, not universal limits.
The script should also state information that viewers need but could miss in the visuals. For example, say “Select the billing settings option” instead of “Select the green button.” The W3C guidance on audio and video recommends clear speech, useful pauses, less jargon, and narration that includes essential visual information.
Record the narration early enough for the animation to follow it. Natural pauses and emphasis affect scene length. If the animation is already locked, the editor may need to rush the narration or remove useful visual explanations.
Audition the actual script
A demo reel shows past work. A custom audition is a short recording of the narrator reading your script. It shows whether the narrator can pronounce your terms, handle the sentences, and match the intended tone.
Choose a representative excerpt of about 20 to 40 seconds. Include a central product benefit, an unusual term, a transition, and a key action or emotional line. Give every candidate identical material and direction. If the product name is unfamiliar, provide a pronunciation guide.
Compare the recordings with a storyboard, animatic, or rough cut instead of listening to them alone. Check whether the voice leaves enough space for visual information and stays clear when music plays. Production guidance for explainer narration also recommends script-specific auditions and recording narration before the animation is finished.
Do not cast by gender, age, accent, or vocal depth alone. One study of 202 U.S. participants found differences in trust ratings among four particular voice treatments, but that result does not establish a universal preference. Test the same script with the audience and context that matter to your project.
Choose human or synthetic narration by project needs
A human voice actor is usually the safer choice when the video depends on emotional nuance, humor, sensitive subject matter, persuasive storytelling, or a distinctive brand personality. Live direction also makes it easier to adjust emphasis and phrasing during the recording session.
Modern synthetic narration fits straightforward copy, frequent product updates, multiple versions, or tight production schedules. Research on older machine voices found advantages for human narration. Later studies using improved text-to-speech systems found similar results for modern synthetic and recorded human voices in some instructional settings. For example, one 2019 study found no statistical difference on most measured learning and perception outcomes, while a 2022 study with 51 undergraduates found no significant difference between modern synthetic and human narration for learning outcomes or cognitive load. These findings come from educational settings, so they do not settle questions about brand trust or commercial persuasion.
Test the complete mix before deciding. A short synthetic sample may sound convincing, yet longer passages can reveal repetitive emphasis, awkward pauses, or incorrect pronunciation. If you use a digital replica of a real person’s voice, obtain explicit written consent and define the permitted uses. The U.S. Copyright Office’s report on digital replicas describes agreements that limit voice use by purpose, location, and duration.
Record clean audio and mix for speech
A good performance can still fail if the recording contains echo, fan noise, or microphone pops. Plosives are bursts of air from sounds such as “p” and “b.” A pop filter helps reduce them. Room reflections happen when sound bounces off hard surfaces, making speech sound distant or echoey.
Record in a quiet, non-echoing space with a suitable microphone. Ask for clean, minimally processed files when an editor will complete the final mix. Heavy compression, which reduces the difference between loud and quiet moments, and aggressive EQ, which changes frequency ranges, can limit later adjustments.
Keep music below the narration. W3C guidance for the relevant accessibility level recommends background audio at least 20 decibels below foreground speech, except for very short sounds. Check the final video on headphones and ordinary laptop or phone speakers, because words that sound clear in a studio can disappear on small speakers.
Plan captions and rights before release
Create accurate captions and review them manually. Automatic captions can miss words, punctuation, speaker changes, and important sound information. Section 508 guidance distinguishes captions, which synchronize text with audio, from transcripts, which provide the words without timing.
Before recording, put the business terms in writing. Confirm whether the fee covers paid advertising, social cutdowns, translations, revisions, territories, duration, and future campaigns. Clarify who receives the final recording and whether raw takes are included. A licensing agreement typically defines the medium and length of use, as described in voice-over contract guidance. Ownership and work-made-for-hire rules depend on the facts, so substantial campaigns deserve legal review.
Final approval checklist
Before publishing, confirm that:
- the voice fits the audience and purpose;
- the script sounds natural when read aloud;
- product names and technical terms are pronounced correctly;
- the pace leaves room for the visuals;
- the audition was tested against the rough cut;
- the recording has no distracting noise or echo;
- music does not mask important words;
- captions, plus a transcript when useful, have been reviewed;
- essential visual information appears in the narration;
- usage rights cover the actual distribution plan.