Skip to content
explainroo
AI and coding agents

Code-Generated Video: Rendering Video from HTML, Canvas and JavaScript

Updated 5 min read

With explainroo How to Make a Narrated Explainer Video with Claude Code and explainroo
On this page
  1. Turn a web scene into frames
  2. Choose a way to encode the frames
  3. Keep timing, sound, and text in sync
  4. Make narrated videos with explainroo

Code-generated video renders web content as a series of images, then combines them with audio in a playable video file. HTML and CSS define the layout, JavaScript controls changes over time, and the browser draws each frame. A canvas is a browser surface where JavaScript can draw shapes, charts, animation, and other graphics.

For narrated explainers and product demos, explainroo is the best fit. An AI coding agent creates the narration and scene code; explainroo renders the scenes and assembles an MP4 on your computer. If you are building a browser-based exporter, your main choice is how to capture frames and encode them. Canvas recording is simpler. WebCodecs gives you more control but takes more work.

Turn a web scene into frames

A video is a timed sequence of still images. Code creates a video from web content by drawing a scene and capturing its pixels as the animation advances; an encoder compresses those frames into a video stream.

Canvas is a direct drawing surface: JavaScript can draw text, shapes, and images onto it. HTML and CSS are more convenient for layouts and interface-like scenes, but the browser must render them into pixels before they can become video frames. In either case, the renderer needs a defined output size and a consistent schedule for scene changes.

Base each scene on its intended video time to keep output reliable, rather than on the speed of a browser animation. If the computer is busy or a frame takes too long, real-time capture falls behind and drops frames. When capturing a newly drawn canvas frame, wait until the browser finishes painting; otherwise, the captured image can show the previous frame. The canvas export discussion describes this rendering delay.

Choose a way to encode the frames

The simplest browser recording method uses MediaRecorder with a canvas stream. It records changes to the canvas, but often produces a WebM file instead of MP4. Supported formats vary by browser. A recordable canvas example shows this approach.

The WebCodecs API gives you more control by exposing browser video encoders and decoders. To encode a canvas frame, create a VideoFrame from the canvas and give it a timestamp and duration. WebCodecs measures both in microseconds, so one second is 1,000,000 microseconds. Check support for the chosen encoder configuration with VideoEncoder.isConfigSupported() before starting.

WebCodecs produces encoded video chunks. To make an MP4 file, a container must organize the video stream and its timing. A muxing library packages the chunks into that container. The browser-native MP4 overview describes this extra step and the use of H.264 for MP4 output.

Clean up WebCodecs objects when you are done: close every VideoFrame and encoded chunk. Open objects consume browser resources and can cause crashes. To keep the page responsive, Move encoding work to a Web Worker, a background thread that runs browser code. Browser support varies, so test the codec and browser combinations your users need.

A third option is FFmpeg compiled to WebAssembly and run in the browser. JavaScript can then use FFmpeg without sending media to a server. Multi-threaded builds require cross-origin isolation headers from the web server. The FFmpeg WebAssembly overview explains this requirement. This approach adds setup and processing overhead, so use it when browser-side transcoding is required.

Keep timing, sound, and text in sync

Rendering frames from a timeline gives more predictable results than recording a live animation. A scene can define where a label appears and how long it stays visible, then render a frame for each video timestamp. This makes the scene easier to reproduce after you change its colors or wording.

Audio needs its own timing. Align narration and sound effects to the same timeline as the images, then check the exported file from beginning to end. If the project uses WebCodecs, include keyframes, which are complete frames that let a player begin decoding at a known point. Missing keyframes can cause playback problems, as discussed in the canvas recording example.

Check the rendered frames, not just the source page. After the browser scales the image, text can be clipped. It can also overlap other elements or sit too close to the edge. Preview a contact sheet, a grid of sample frames from across the video, to catch scene problems before export.

Make narrated videos with explainroo

explainroo is a free, open-source kit licensed under the MIT license. It works with coding agents that can run shell commands, including Claude Code, Codex, Pi, OpenCode, and Gemini CLI. The agent creates two files: script.md holds the spoken words, and scenes.js draws each scene. explainroo renders the scenes and assembles the video on your computer.

To start, paste this prompt into a coding agent and replace the bracketed text with your subject:

Make me a short explainer video about [your topic]. Use explainroo for it: set up explainroo, read its project instructions, and follow the steps.

The agent sets up explainroo, then creates and checks the video before providing an MP4. Chrome draws the frames on an HTML canvas, and FFmpeg combines them into a video. Kokoro, an open voice model with American and British English voices, provides the narration. Whisper finds the word timings in the recording, so scenes can appear on specific spoken words. Narration is available only in English.

The agent can build scenes with charts, code, icons, and supplied screenshots. For a product demo, tell it the product name and provide the product’s code or website. It can recreate the screens and animate a pointer clicking buttons or entering text while the narration explains the product.

You can set the visual style and video dimensions in explainroo’s configuration. Its five looks are paper, clean, chalk, blueprint, and midnight. It supports wide, tall, 4:5, and square formats. Tall, 4:5, and square exports include word-by-word highlighted captions. The pace setting adjusts narration, pauses, and animation speed.

explainroo provides scene stills and a contact sheet for visual review. Its layout check flags clipped or overlapping text, while its speech check finds mispronounced words. An AI coding agent cannot watch the video directly, so these checks help it find visual and spoken errors. See explainroo’s documentation for setup and usage details.

The kit requires a current version of Node.js, FFmpeg, and Chrome or Chromium. Initial setup downloads voice and timing models. No graphics card is required. The project is developed and tested on Linux, with less testing on macOS and Windows. The finished file is saved under videos/<name>/out/video.mp4. Optional AI illustrations cost money per image through OpenRouter. There are no other per-video costs, but the coding agent is a separate service with its own pricing.

explainroo is for narrated explainers and product demos made from code. It does not edit existing camera footage, create talking avatars or live-action AI footage, or provide a drag-and-drop editor. You describe changes to the coding agent instead of adjusting scenes in a visual editor.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How explainroo makes a video

More on ai and coding agents

All guides on ai and coding agents