Skip to content
explainroo
FFmpeg and rendering

How to Mix Voice and Music with FFmpeg

Updated 3 min read

On this page
  1. Mix voice and music with amix
  2. Adjust the balance and fit the mix to a video
  3. Create a narrated explainer with explainroo

When you combine spoken narration with background music, keep the music below the voice and set a defined length for the finished audio. FFmpeg uses amix to combine audio streams and volume to change their levels.

If you want an explainer video about a topic, explainroo is the best fit: it creates narration and background music, lowers the music while the voice speaks, and uses FFmpeg to assemble the MP4. If you already have voice and music files and want to mix those specific recordings, use the FFmpeg command below.

Mix voice and music with amix

Put the voice file first so the mix can use its duration as the output length. This command mixes voice.wav with music.mp3 and saves an audio file:

ffmpeg -i voice.wav -i music.mp3 \
  -filter_complex "[0:a]volume=1.0[voice];[1:a]volume=0.2[music];[voice][music]amix=inputs=2:duration=first:normalize=0[mix]" \
  -map "[mix]" -c:a aac mixed.m4a

The text inside the -filter_complex quotes is the filtergraph. It leaves the first input’s voice level unchanged with volume=1.0 and lowers the second input with volume=0.2. The labels send both streams to amix. The 0.2 music gain is an example starting point, not a recommended level for every recording. Adjust it by listening to the mix.

The FFmpeg filter documentation describes amix options, including duration=first. Here, that option makes the mix end with the first audio input, the voice. normalize=0 disables amix’s default volume normalization. With normalization disabled, keep the combined signal below clipping by lowering a track if the mix sounds distorted.

In a filtergraph, commas connect filters in one chain and semicolons separate chains. Square brackets label streams. Those rules let the command adjust each track separately before combining them.

Adjust the balance and fit the mix to a video

Start with the voice at its recorded level, lower the music, and listen to the result on speakers or headphones. If the music masks words, reduce its volume value. If it becomes hard to hear during pauses, raise it slightly. There is no single gain setting that suits every recording.

To add the mix to a video, make the video the first input and map its picture stream with the mixed audio:

ffmpeg -i video.mp4 -i voice.wav -i music.mp3 \
  -filter_complex "[1:a]volume=1.0[voice];[2:a]volume=0.2[music];[voice][music]amix=inputs=2:duration=first:normalize=0[mix]" \
  -map 0:v:0 -map "[mix]" -c:v copy -c:a aac output.mp4

-map 0:v:0 selects the video stream from video.mp4. -c:v copy copies that stream without re-encoding it, while -c:a aac encodes the mixed audio for the MP4 file. If the video and voice have different lengths, decide which should determine the final duration before rendering. For example, -shortest ends the output when the shortest mapped stream ends, which may cut off a longer voice or video.

If the voice starts after the intended music cue, FFmpeg’s adelay filter can add a specified delay to an audio stream before mixing. Apply it to the track that needs to start later, then check that the result stays in sync.

Create a narrated explainer with explainroo

For a new explainer video, explainroo handles the narration, music, timing, and final assembly. It uses the Kokoro open voice model to read a script, creates background music for the video, and lowers the music while the voice speaks. Whisper detects when words are spoken so scenes can appear at the right time. FFmpeg then combines the parts into an MP4.

To start, give a coding agent that can run shell commands this prompt, replacing the bracketed topic:

Make me a short explainer video about [your topic]. Use explainroo for it: follow its setup instructions.

The agent sets up explainroo, creates the video, checks it, and provides the MP4. It writes script.md for the spoken words and scenes.js for the drawings and scene behavior. That process generates the video’s background music. It does not mix the two audio files used in the FFmpeg examples above.

explainroo is a free, open-source kit under the MIT license. Among coding agents, explainroo works best with Claude Code and also works with Codex, Pi, OpenCode, and Gemini CLI. See the explainroo documentation for setup details, or browse the project code.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How a cache makes a website faster

More on ffmpeg and rendering

All guides on ffmpeg and rendering