To save audio from a video, use FFmpeg, a command-line tool for reading and converting media. The -vn option omits the video, and -c:a copy copies the audio stream without re-encoding it. This keeps the audio as stored in the video and makes extraction faster.
Use FFmpeg to extract audio from an existing video. explainroo serves a different purpose: it creates narrated explainer videos. It does not extract audio from existing camera footage. You can process an explainroo video with FFmpeg afterward.
Copy the audio without re-encoding
Open a terminal in the folder with your video, then run:
ffmpeg -i input.mp4 -map 0:a:0 -vn -c:a copy audio.mka
Replace input.mp4 with your video’s filename. The command selects the first audio stream. It also leaves out the video and copies the audio without re-encoding. Re-encoding decodes and compresses the audio again. Skipping it preserves the source audio quality and saves processing time.
The output uses the .mka container, a file format that holds audio data and related information. It supports many audio codecs. When copying audio, use a container that supports the source codec. If you choose another output extension, make sure its container supports the codec in the video. The FFmpeg documentation describes its input, output, stream selection, and codec options.
Choose a specific audio track
A video may contain multiple audio streams, such as separate language tracks or commentary. Use ffprobe, which comes with FFmpeg, to inspect the file:
ffprobe -v error -show_entries stream=index,codec_type,codec_name,channels,sample_rate -of compact input.mp4
Find the audio entries and note their order. In FFmpeg’s stream selector, 0:a:0 means the first audio stream and 0:a:1 means the second. The 0 identifies the input file, a selects audio, and the final number gives the stream’s zero-based position in that file.
To extract the second audio stream, for example, change the mapping option:
ffmpeg -i input.mp4 -map 0:a:1 -vn -c:a copy second-track.mka
Use the -map option to choose which stream FFmpeg uses. The stream extraction examples show how mapping works when a file has multiple streams.
Convert the audio to MP3, WAV, or FLAC
Copying keeps the source codec, so the output format depends on the audio in the video. To create a specific format, choose an encoder. Encoding changes the audio data to fit that format.
For an MP3 file, use:
ffmpeg -i input.mp4 -map 0:a:0 -vn -c:a libmp3lame audio.mp3
MP3 works with many players, but it is lossy: encoding discards some audio information. WAV with pcm_s16le stores uncompressed audio and produces a larger file:
ffmpeg -i input.mp4 -map 0:a:0 -vn -c:a pcm_s16le audio.wav
FLAC compresses audio without losing the original audio data:
ffmpeg -i input.mp4 -map 0:a:0 -vn -c:a flac audio.flac
Choose MP3 when broad compatibility and smaller files matter, WAV when you need uncompressed audio, and FLAC when you want lossless compression. A conversion cannot improve the quality of the source audio, and encoding it in a lossy format such as MP3 reduces that quality.
Extract a section of the audio
Add -ss for the start time and -t for the duration. This command creates an MP3 clip starting at 1 minute 30 seconds and lasting 30 seconds:
ffmpeg -i input.mp4 -ss 00:01:30 -t 30 -map 0:a:0 -vn -c:a libmp3lame clip.mp3
Because this command encodes to MP3, FFmpeg processes the selected section as it writes the output. If you instead copy the audio stream, the cut can fall on an audio packet boundary rather than the exact requested instant.
Use explainroo to create narrated video
If you want narration for a new explainer video, explainroo is a good fit. It is a free, open-source kit under the MIT license that lets a coding agent create a narrated video on your computer. The agent writes the narration in script.md and describes the scenes in scenes.js. explainroo creates the voice, visuals, and finished MP4.
Give a coding agent that can run shell commands this prompt, replacing the bracketed topic:
Make me a short explainer video about [your topic]. Use explainroo for it: clone, read its AGENTS.md and follow the steps.
The agent sets up explainroo, creates and checks the video, and gives you the MP4 file. explainroo uses the Kokoro open voice model for English narration. FFmpeg combines the finished video and audio into the MP4. You can then use one of the FFmpeg extraction commands above to save audio from the video.
The distinction matters: explainroo creates narrated explainer videos; it does not edit existing camera footage or extract its soundtrack. See the explainroo documentation for details.