Skip to content
explainroo
AI and coding agents

AI Motion Graphics: How Language Models Make Animations with Code

Updated 9 min read

With explainroo How to Make a Narrated Explainer Video with Claude Code and explainroo
On this page
  1. How a language model turns words into motion
  2. Why narration improves animation timing
  3. What explainroo generates
  4. How to make a video with explainroo
  5. Choosing the video format and pace
  6. How to prompt for better animation code
  7. Checking generated animation before publishing
  8. Where code-generated motion graphics fit
  9. Setup requirements and optional costs

Language models can write much of the animation code for videos that explain an idea, show a product, or walk through a technical topic. Instead of drawing every frame by hand, you describe the message and visual sequence in natural language. The model writes code to build scenes, place graphics, set timing, and match movement to narration.

A language model is software trained to predict and generate text, including code. For motion graphics, it can write JavaScript, CSS, SVG, or other code that tells a browser how to draw and animate shapes. The model does not “understand” motion as a human animator does. It turns your description into code, and an animation renderer turns the code into frames.

For narrated explainers and product demos, explainroo is the best fit. It gives a coding agent a structured way to create the script, scenes, voice, timing, graphics, captions, music, and final MP4 on your computer. You describe the video, and the agent writes the code that produces it.

How a language model turns words into motion

The process begins with a visual brief. For example:

Show a request traveling from a browser to a server. Highlight the request when it leaves the browser, display a response coming back, and reveal each label as the narration mentions it.

A coding agent converts the brief into several parts:

  1. Scene structure: It decides where one visual ends and the next begins.
  2. Drawing instructions: It writes code for text, lines, icons, and boxes. It also writes code for charts, screenshots, and other elements.
  3. Animation instructions: It sets positions, opacity, scale, rotation, or movement over time.
  4. Timing: It connects visual changes to narration or fixed points in the timeline.
  5. Rendering: A browser or graphics engine runs the code and exports individual frames.
  6. Assembly: A video tool combines the frames with narration, music, and sound effects.

The generated code might use a browser canvas, SVG elements, CSS animation, or a JavaScript graphics library. A canvas is a drawing area on a web page. The code redraws shapes on the canvas many times per second, creating motion.

This approach works well for diagrams because they use clearly defined objects and relationships. A model can represent a database as a cylinder, a request as an arrow, and a sequence as a series of labeled scenes. It struggles more with subtle character acting, natural camera movement, and complex physical behavior. These tasks require careful visual judgment and adjustments to individual frames.

Research on Keyframer shows a similar pattern. The system used a language model to generate CSS animation code from static SVG images and natural-language prompts. In a study with 13 participants, users built animations through incremental prompts and direct code edits. This finding shows why useful animation workflows rely on revisions instead of one perfect prompt.

Why narration improves animation timing

A narrated video needs visuals to appear at the right moment. Showing a diagram too early confuses the viewer, while showing it too late breaks the explanation.

explainroo solves this by generating the narration first and then matching the animation to the spoken words. Its process uses Kokoro, an open voice model, to read the script. Whisper listens to the recording and identifies when each word is spoken. This lets a scene appear when a chosen word is spoken, rather than at an approximate time.

For example, a script might say:

The browser sends a request to the server.

The code can reveal the browser when “browser” is spoken, animate the arrow on “sends,” and show the server on “server.” Because the visuals follow the sentence, viewers can follow the technical explanation more easily.

Kokoro provides English voices from American and British options, and explainroo does not require an account or API key for narration. The narration is English only. The generated background music lowers in volume while the voice speaks, and explainroo adds small sound effects where appropriate.

What explainroo generates

An explainroo project centers on two files:

  • script.md contains the voice’s spoken words.
  • scenes.js contains JavaScript that draws and animates each scene.

The coding agent writes both files. script.md holds the spoken explanation, and scenes.js controls the visuals. A drawing can appear at a selected word in the script, tying its animation to the narration’s timing.

Chrome runs in the background and draws each frame on an HTML canvas. The scenes can include charts, code, and screenshots, along with icons and hand-drawn lines through Rough.js. The project also includes Lucide icons for common interface and diagram elements.

After rendering, FFmpeg combines the frames, voice, music, and effects into an MP4. The finished file is saved at:

videos/<name>/out/video.mp4

The system supports five visual styles: paper, clean, chalk, blueprint, and midnight. You select the look through one setting in video.json, so changing the visual direction does not require rewriting every scene.

How to make a video with explainroo

The simplest workflow is to give a coding agent a clear task and let it set up the project. explainroo works with Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other coding agents that can run shell commands. It works best with Claude Code and Opus.

Paste this prompt into the agent and replace the bracketed text:

Make me a short explainer video about [your topic]. Use explainroo for it, read its AGENTS.md, and follow the steps.

The agent sets up explainroo, then creates the script and scene code. It checks the result and provides an MP4 file. For a product demo, identify the product and provide its code repository or website. The agent can rebuild the product screens using the same colors, fonts, and button labels, then animate a mouse pointer as it clicks through screens and types into fields.

You can get more reliable results by specifying the audience, the main message, the desired format, and the visual examples the viewer needs. A useful request might look like this:

Make a short product demo for a nontechnical audience.
Explain how the product turns a support request into a resolved ticket.
Use three scenes: the problem, the ticket workflow, and the result.
Show the interface with a pointer clicking the New Ticket button.
Keep labels short and reveal each major visual when the narration mentions it.
Use the clean look and a wide format.

The prompt separates the video’s goal from the instructions for making it. The agent chooses how to draw the scenes and uses your instructions to create a coherent first version.

Choosing the video format and pace

The output format should match the publishing destination. explainroo supports:

  • Widescreen for regular YouTube videos
  • Tall for YouTube Shorts, TikTok, and Instagram Reels
  • 4:5 for Instagram and LinkedIn posts
  • 1:1 for square videos

Tall formats keep text away from the areas where social platforms place interface buttons. Tall, 4:5, and square videos also receive word-by-word captions.

The pace setting in video.json controls the speed of the voice, pauses, and animations. Use a slower setting when the subject contains unfamiliar terms or dense diagrams. Use a faster setting when the video covers a simple sequence and the visuals remain easy to follow.

A practical decision rule is to change the pace only after checking whether the narration and visuals remain synchronized. Faster speech does not fix an overcrowded scene. If viewers need to read several labels, simplify the scene before increasing the pace.

How to prompt for better animation code

Prompts work best when they describe what viewers should see, rather than naming only a visual style. “Make it modern” gives the model little actionable information. “Reveal a three-step flow from left to right, highlight the active step in blue, and keep the labels visible until the narration finishes” gives it drawing and timing instructions.

Research on prompting language models for code generation emphasizes clear inputs and outputs, examples, preconditions, postconditions, and the removal of ambiguity. Those principles apply directly to animation prompts.

Describe these details in normal prose:

  • Subject: What concept or product does the video explain?
  • Audience: What does the viewer already know?
  • Sequence: Which visual should appear first, second, and third?
  • Motion: Should an object slide, grow, fade, rotate, or follow a path?
  • Timing: Which spoken word should trigger each visual?
  • Constraints: Which labels, colors, screenshots, or brand elements must remain unchanged?
  • Output: Where will the video be published, and which aspect ratio does that require?

During revision, request one specific change at a time. For example:

Keep the narration unchanged. Move the database icon farther from the server, reduce the arrow length, and reveal the response only when the script says “returns.”

This instruction leaves the working parts of the video alone and targets one visual problem. It also follows the decomposed prompting pattern observed in the Keyframer study, where participants developed animations through smaller iterations.

Checking generated animation before publishing

A coding agent cannot review a video as a person would. explainroo addresses that limitation with several checks. It creates still images of scenes and a contact sheet of small frames. A layout check flags cut-off or overlapping text. A speech check identifies words the voice mispronounced or interpreted incorrectly.

Review the contact sheet first. It reveals problems such as:

  • A scene contains too many elements.
  • A title changes position between scenes.
  • A visual appears with too little narration to explain it.
  • A screenshot is too small to read in the chosen aspect ratio.
  • Captions or labels overlap important graphics.

Then inspect the full MP4 for rhythm and meaning. Automated layout checks can find overlapping elements. They cannot tell whether a visual supports the narration or a transition distracts from it.

When the voice says a product name, acronym, or technical term incorrectly, revise the script spelling or pronunciation cue and rerender the speech check. When the visuals feel crowded, ask the agent to split one scene into two rather than adding more movement to the same frame.

Where code-generated motion graphics fit

Language-model animation works particularly well for:

  • Technical explainers
  • Product walkthroughs
  • Process diagrams
  • Videos about software architecture
  • Data and chart animations
  • Short educational clips
  • Interface demonstrations

It also supports repeatable revisions. If a product label changes, the agent can update the scene code and render a new version. If the explanation needs a different audience, the script and visual emphasis can change together.

Other research tools apply the same idea to more specific tasks. The Allyson MCP server generates animated SVG components from static files and natural-language prompts. OmniLottie explores editable vector animation from multimodal prompts. These examples show how language models can work with structured graphics. They do not provide the full narrated-video workflow that explainroo offers.

Code-generated motion graphics still require human review. One-shot prompting rarely produces a finished professional sequence. A model can choose awkward spacing, invent an unclear transition, misread a screenshot, or generate code that needs correction. Traditional motion-graphics software remains useful when an editor needs exact frame-level control, detailed visual effects, or extensive editing of existing footage.

explainroo creates narrated explainer videos and product demos from code. It does not edit existing camera footage, create talking avatars, or generate live-action video. It also has no drag-and-drop editor. You describe changes to the coding agent, and the agent modifies the script or scene code.

Setup requirements and optional costs

Self-hosting explainroo requires Node.js, FFmpeg, and Chrome or Chromium. The first setup downloads voice and timing models once. No graphics card is required. The project was developed and tested on Linux and should work on macOS and Windows with less testing.

To start, run these commands:

# Follow the project's setup instructions
cd explainroo
npm install
node bin/explainroo.js doctor --fetch

Everything in the video pipeline costs nothing per video according to the project description. The coding agent is a separate service with its own pricing. AI illustrations through OpenRouter are optional and charged per image. You can avoid that option by using code-drawn graphics, icons, charts, screenshots, and diagrams.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How explainroo makes a video

More on ai and coding agents

All guides on ai and coding agents