Skip to content
explainroo
AI and coding agents

How to Make Videos with Gemini CLI

Updated 7 min read Published by Vincent Schmalbach, written with Unspar

On this page
  1. What Gemini CLI can do for a video project
  2. Install and start Gemini CLI
  3. Give Gemini CLI project instructions
  4. Generate the script and shot list
  5. Add images, recordings, and generated video clips
  6. Use Gemini CLI to assemble the project
  7. Automate repeatable video jobs
  8. Troubleshoot the limits

Gemini CLI is Google’s open-source AI agent for the terminal, a text-based tool for running commands and managing files. It can inspect a project and write or edit files. It also runs shell commands and uses Google Search to find supporting information.

Gemini CLI has no documented built-in command that turns a text prompt into a finished video. Use it as a production assistant instead: generate the script and shot list, create or organize assets with external tools, and assemble the final video with a local video encoder or an external video service.

This distinction matters because Google’s video-generation page describes video creation in the Gemini app. The available Gemini CLI documentation focuses on terminal-based coding, file work, command execution, and tool extensions. Gemini CLI can support a video workflow, but direct Veo or Gemini video generation from the CLI requires a separately configured integration.

What Gemini CLI can do for a video project

Gemini CLI handles video-production tasks that use text, files, and repeatable commands. For example, it can:

  • Write a script, narration, captions, and scene descriptions
  • Turn a rough idea into a shot list
  • Create project folders and text files
  • Rename and organize media files
  • Inspect existing assets and metadata
  • Generate commands for a local video tool
  • Run approved shell commands
  • Connect to external services through MCP servers or CLI extensions

An MCP server connects an AI client to an external tool through a defined interface. An extension adds commands or other capabilities to Gemini CLI. The Google Cloud guide to Gemini CLI describes both approaches.

Gemini CLI can plan and automate many production tasks, while lacking a video-generation model and a complete video editor.

Install and start Gemini CLI

Gemini CLI runs through Node.js and is available through npm, Homebrew, or npx. Guides list different required Node.js versions, so check the current official Gemini CLI site before installing. An outdated Node.js version commonly causes installation failures.

After installation, start the command-line client from a terminal:

gemini

Sign in with a personal Google account when prompted. The free tier does not require an API key or credit card, according to the Gemini CLI installation guide. It has usage limits that may change, so check the current account terms before relying on it for batch production.

Create a separate directory for each video:

mkdir product-demo
cd product-demo
mkdir -p script assets audio renders

These folders keep planning files separate from source media and exported videos. This makes it easier to repeat a render or hand the project to another editor.

Give Gemini CLI project instructions

Create a GEMINI.md file in the project directory. Gemini CLI loads this file as persistent project context during sessions. Use it to define the video’s audience, tone, format, naming rules, and production constraints.

For example:

## Video project instructions

Create short educational videos for technically comfortable beginners.

Use:

- Clear US English
- Short sentences
- One idea per scene
- Time codes in minutes and seconds
- Separate narration from on-screen text
- File names in lowercase with hyphens

Do not invent product capabilities or performance claims.

The file gives Gemini CLI stable instructions, so you do not need to repeat them in every prompt. Keep factual constraints here, especially when the video describes a product, medical topic, financial service, or regulated claim.

Generate the script and shot list

Ask Gemini CLI to create structured production files instead of one large block of prose. Start with a prompt such as:

Create a 60-second explainer video package for this topic:

[describe the topic]

Write these files:
1. script/script.md with narration and time ranges
2. script/shot-list.md with one row per shot
3. script/on-screen-text.md with captions and titles

For each shot, include:
- approximate duration
- visual description
- narration
- on-screen text
- required asset filename

Keep the narration within the requested duration. Mark any claim that needs fact-checking.

The shot list should connect every spoken line to a visual. A product demonstration, for example, might include a screen recording, a close-up of the relevant interface, and a final title card. This prevents a common editing problem: finishing the voice-over before finding matching visuals.

Ask Gemini CLI to review the draft before producing assets:

Review the script and shot list for:
- narration that exceeds the planned duration
- repeated ideas
- unsupported factual claims
- on-screen text that is too long to read
- shots with no available asset

Write the findings to script/review.md.

Gemini CLI can read and edit project files, so the review can happen in the same directory as the script.

Add images, recordings, and generated video clips

Gemini CLI does not create a video clip by itself through a documented built-in video command. You need to supply footage from one or more external sources:

  • Screen recordings
  • Camera footage
  • Images and illustrations
  • Stock media
  • Clips produced by a separate video-generation service
  • Audio narration from a voice tool

Store the source files in assets/ and use filenames that match the shot list. For example:

assets/
  opening-screen.png
  product-demo.mp4
  narrator.wav
  background-music.wav

If you want to use an external video-generation provider, check whether it offers an API, command-line interface, MCP server, or Gemini CLI extension. Gemini CLI needs a connector before it can call that provider. The available documentation does not establish a direct Gemini CLI connection to Veo, so a prompt inside Gemini CLI will not necessarily generate a Veo clip.

You can still ask Gemini CLI to prepare prompts for a separate generator:

Using script/shot-list.md, create script/video-prompts.md.

For each shot that needs generated footage, provide:
- subject
- action
- camera movement
- framing
- lighting
- duration target
- continuity details

Keep each prompt self-contained and do not add text overlays inside generated footage.

Review the prompts manually before sending them to another service. Video models frequently interpret character appearance, camera movement, and text differently from the wording you provide.

Use Gemini CLI to assemble the project

After you have the assets, Gemini CLI can help create an assembly plan and run approved local commands. First, ask it to inspect the project:

Inspect the files in assets/ and script/shot-list.md.

Identify:
- missing assets
- mismatched filenames
- clips that do not match the planned shot
- audio files that need conversion
- the intended scene order

Write the result to script/assembly-notes.md. Do not modify media files yet.

For rendering, use a local video tool that you have installed and understand. FFmpeg is one common command-line encoder. The exact command depends on the source formats and frame rates. It also depends on the audio tracks, transitions, subtitles, and output requirements.

Have Gemini CLI draft the command before execution:

Based on script/shot-list.md and script/assembly-notes.md, draft the local video-rendering commands.

Requirements:
- preserve the intended scene order
- include narration
- include captions if the files support them
- write output to renders/final.mp4
- do not delete or overwrite source files

Show the commands for review. Do not run them yet.

Review every command that reads, moves, overwrites, or deletes files. Gemini CLI has approval modes that range from normal confirmation to automatic editing and fully automatic command execution. Its headless mode, invoked with -p, supports scripting and piping. Automatic modes suit controlled batch jobs, but running commands without approval is risky in a directory that contains valuable files or untrusted content.

Automate repeatable video jobs

After one project works manually, move repeated tasks into a script. For example, a shell script could:

  1. Check that required assets exist.
  2. Ask Gemini CLI to update the shot list.
  3. Run the local render command.
  4. Put the result in a dated output directory.
  5. Save a render log.

Gemini CLI’s headless mode is designed for noninteractive use, so it fits scripted workflows. Keep the first version conservative. Require the script to stop when an asset is missing, and write to a new output filename instead of replacing the previous export.

A GEMINI.md file can also define stable rules for batch work, such as:

Before rendering:

- Confirm every asset referenced by the shot list exists.
- Never delete source media.
- Never overwrite files in assets/.
- Stop when narration and video durations do not align.
- Save generated outputs under renders/.

These instructions reduce accidental changes, but they do not replace command review or file backups.

Troubleshoot the limits

When Gemini CLI cannot make a video, check which part of the workflow is failing.

If it writes a script but produces no moving footage, that is expected without an external generation or editing tool. If it proposes a Veo command that does not work, verify that you have an official API or configured extension rather than assuming the Gemini app and Gemini CLI expose the same capabilities. If installation fails, check the current Node.js requirement because published guides have listed different minimum versions.

Free-tier limits also affect batch jobs. The reported allowance includes per-minute and daily request limits, but those values are subject to change. Design long workflows to resume from saved files instead of regenerating the entire project after one failed request.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How explainroo makes a video

More on ai and coding agents

All guides on ai and coding agents