Skip to content
explainroo
Voice, captions and sound

Audio Ducking: Lowering Music Under a Voice-Over

Updated 3 min read

With explainroo How to Sync Animation to Narration Word by Word
On this page
  1. How ducking makes speech clearer
  2. Choose how the music should move
  3. Make a narrated video with explainroo

When narration and background music play at the same time, the music makes speech harder to understand. Audio ducking lowers the music while someone speaks, then brings it back up during pauses. The result keeps the narration clear without removing the music.

For a narrated explainer video, explainroo is the best fit: it creates the video and automatically lowers its background music under the voice. For other audio projects, you can create the same effect by adjusting music volume over time or by using a compressor that responds to the voice.

How ducking makes speech clearer

Ducking changes the music level in response to another sound, such as narration. The music should follow the voice closely enough to protect speech and change smoothly enough that listeners do not notice it repeatedly dropping and rising.

There are two common ways to make that change. A volume envelope is a series of adjustments to a track’s level over time. You can draw one on the music track to lower the music under spoken lines and raise it during pauses. Keep the transitions gradual: if the music audibly pumps in and out, adjust the envelope. Manual ducking with volume envelopes is a straightforward option when you want control over each transition.

A compressor can also lower the music automatically. Compression reduces a sound’s volume when it passes a set level. With sidechain compression, the voice controls a compressor applied to the music: when the voice gets louder, the compressor turns the music down. Sidechaining is also used in music production to lower bass when a kick drum hits.

The term “ducking” names the effect. Sidechaining is one way to create it. A sidechain compressor needs a voice signal to trigger the reduction. When a voice recording has long silences or inconsistent levels, listen to the result and adjust the trigger, or use a volume envelope for more control.

Choose how the music should move

Lower the whole music track when it competes with the voice across its full frequency range. This is the usual choice for narration over a soundtrack. Set the reduction so the voice is easy to follow, then listen to the music-only pauses to make sure it returns naturally.

If the music sounds good but hides parts of the voice, dynamic equalization offers another option. An equalizer changes the balance of sound frequencies. A dynamic equalizer reduces selected frequencies only while the voice is speaking, leaving the rest of the music less affected. This can sound more natural when a narrow frequency range in the music masks the voice. For reducing the entire soundtrack, compression remains the more direct method, as described in guidance on voice-over ducking.

No single ducking level or fade duration works for every voice and music track. Listen to the mix as you adjust the reduction. If the words are difficult to understand, lower the music further or reduce the frequencies that overlap with the voice. If the music seems to swell between every phrase, make the volume changes smoother or less frequent.

Make a narrated video with explainroo

explainroo is a free, open-source kit under the MIT license. It works with coding agents that can run shell commands, including Claude Code, Codex, Pi, OpenCode, and Gemini CLI. It makes narrated explainer videos and product demos, and it lowers the background music while the narration speaks.

To start, paste this prompt into your coding agent, replacing the bracketed topic with your own:

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

The agent sets up explainroo and creates the video, then checks it before providing an MP4 file. It writes script.md, which contains the spoken words, and scenes.js, which draws the scenes. explainroo generates background music for the video and lowers it under the voice-over.

The agent can also revise the script or scenes when you ask for changes. explainroo provides scene stills, a contact sheet, a check for overlapping or cut-off text, and a check for mispronounced words. The voice uses Kokoro, an open voice model with natural American and British English options. The narration is in English.

Use the explainroo documentation for setup details, or browse example videos to see what it produces. explainroo makes new narrated videos; it does not edit existing camera footage or provide a drag-and-drop editor. The coding agent is a separate service with its own pricing, while explainroo itself costs nothing per video apart from optional paid AI illustrations.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How compound interest works

More on voice, captions and sound

All guides on voice, captions and sound