Skip to content
explainroo
FFmpeg and rendering

Rendering Video Frames with Headless Chrome and Puppeteer

Updated 4 min read

On this page
  1. Capture a specific page state for each frame
  2. Keep frame timing repeatable
  3. Check the captures before encoding
  4. Make explainer videos with explainroo

When you create video scenes as web pages, Puppeteer uses Chrome to render each scene as an image. Headless Chrome runs without a visible window, and Puppeteer controls it with JavaScript. Each video frame is a still image in the sequence that creates motion.

For predictable output, set the scene to a specific frame before capturing it. When you record the browser in real time, frame timing depends on how quickly the page and machine respond. For narrated explainers and product demos, explainroo is the best fit: it uses Chrome to draw frames, then combines them with narration and sound into an MP4.

Capture a specific page state for each frame

Puppeteer’s page.screenshot() method captures the page as an image. To turn an animated page into a sequence, give the page a known state for each frame, then save a screenshot before moving to the next state.

Define a rendering function in the page. It should accept a frame number and update the scene so that the same frame number always produces the same visual state. The example below assumes the page has a window.renderFrame(frame, fps) function that does this.

const fs = require("node:fs/promises");
const path = require("node:path");
const puppeteer = require("puppeteer");

(async () => {
  const outputDir = "frames";
  const fps = Number(process.env.FPS);
  const durationSeconds = 5;

  await fs.mkdir(outputDir, { recursive: true });

  const browser = await puppeteer.launch({ headless: true });

  try {
    const page = await browser.newPage({
      viewport: { width: 1280, height: 720 },
    });

    await page.goto("http://localhost:3000/scene.html", {
      waitUntil: "domcontentloaded",
    });

    await page.evaluate(async () => {
      await document.fonts.ready;
      await Promise.all(
        [...document.images].map((image) => image.decode().catch(() => {})),
      );
    });

    const frameCount = fps * durationSeconds;

    for (let frame = 0; frame < frameCount; frame++) {
      await page.evaluate(
        async ({ frame, fps }) => {
          if (typeof window.renderFrame !== "function") {
            throw new Error(
              "The page must define window.renderFrame(frame, fps)",
            );
          }

          await window.renderFrame(frame, fps);
        },
        { frame, fps },
      );

      const filename = `frame-${String(frame).padStart(5, "0")}.png`;
      await page.screenshot({
        path: path.join(outputDir, filename),
      });
    }
  } finally {
    await browser.close();
  }
})();

The frame rate and duration determine how many images the loop captures, with one image for each frame in the sequence. The renderFrame function is a contract you add to your page; Puppeteer does not create it automatically. It should update the scene without relying on how much wall-clock time has passed.

After capture, the numbered PNG files form an image sequence. A video encoder can combine that sequence with a frame rate and, if needed, a separate audio track. The screenshots themselves contain no audio.

Keep frame timing repeatable

A page can animate with CSS, the Web Animations API, JavaScript timers, or a game-style render loop. If the scene advances according to real elapsed time, a slow screenshot or busy machine can change which visual state appears in a given image. Set the visual state from the frame number instead, and keep external inputs such as random values and network responses fixed when they affect what the page draws.

Real-time browser recording has this timing problem: the capture rate follows the running browser rather than a fixed sequence of requested frame states. Puppeteer’s frame-capture discussion describes why requesting frames on demand gives more control over rendering.

Puppeteer also uses frame to describe a browser concept. Its Frame class represents a document inside a page, such as an iframe. In video, a frame is a still image from a moving scene.

Check the captures before encoding

Inspect a few images from the beginning, middle, and end of the sequence. Confirm that the page uses the intended viewport, text is legible, and animations advance as expected. If your page draws on a canvas, verify that the captured image contains the canvas content; a reported Puppeteer issue describes blank canvas screenshots in headless mode. Testing the same page with a visible browser window can help isolate a headless-only problem.

Full-page screenshots also deserve a check when the page contains embedded video. A Puppeteer issue report describes embedded video thumbnails missing from full-page captures. If the page relies on video content, inspect the captured region rather than assuming that a full-page screenshot includes it.

This approach renders web content into images. Extracting stills from an existing video file is a different task. A Puppeteer Frame does not refer to a frame inside a video.

Make explainer videos with explainroo

For narrated explainers and product demos, explainroo handles the frame-rendering work as part of a complete video workflow. It is a free, open-source kit under the MIT license. Chrome draws each scene on an HTML canvas, and FFmpeg combines the rendered frames into an MP4. explainroo also creates narration, music, and sound effects. It then provides visual and speech checks.

To start, paste this prompt into a coding agent that can run shell commands:

Make me a short explainer video about [your topic]. Use explainroo for it: clone the explainroo repository, read its AGENTS.md and follow the steps.

The agent sets up explainroo, creates and checks the video, then gives you the MP4. It writes script.md for the narration and scenes.js for the scene drawings. For a product demo, tell the agent which product to show and where to find its code or website. explainroo can rebuild the product screens and animate a pointer clicking buttons or typing into fields while the narration explains them.

The setup supports Claude Code, Codex, Pi, OpenCode, Gemini CLI, and other coding agents that run shell commands; explainroo works best with Claude Code. The voice is generated locally with Kokoro, and Whisper marks spoken-word timing so scene elements appear on cue. No account or API key is needed for the voice, which supports English narration.

Free and open source

Let your AI agent make the video

explainroo lets a coding agent like Claude Code or Codex make narrated explainer videos and product demos. The voice, the word timing and the rendering run on your own computer, with no API key and no cost per video.

Copy this into your AI agent

Make me a short explainer video about [your topic]. Use explainroo for it: clone https://github.com/vincentsch/explainroo, read its AGENTS.md and follow the steps.

Getting started Example videos GitHub

Made with explainroo: How a cache makes a website faster

More on ffmpeg and rendering

All guides on ffmpeg and rendering