Docs
Images
How to use your own screenshots and photos, and how to make pictures with AI image models.
How do I use my own images?
Put the files in the project’s assets folder and draw them with s.image:
s.image("assets/dashboard.png", {
w: 1100,
frame: "browser",
url: "app.example.com",
at: "dashboard",
});
s.image("assets/phone.png", { h: 800, frame: "phone", at: 0.5 });
PNG, JPEG, WebP, GIF and SVG files work. The frames browser, window and phone put a screenshot inside a drawn browser, window or phone. card puts a white border around a picture. kenburns: true adds a slow zoom.
When do AI images help?
Icons and diagrams cover most technical topics. For everyday how-to topics, like cooking, cars or plants, a picture of the food, the car or the plant explains more. explainroo can make these pictures with AI image models through OpenRouter.
This is optional. It costs money on your OpenRouter account, about 7 to 13 cents per image at the time of writing.
How do I set it up?
You need an OpenRouter API key from openrouter.ai/keys. Put it in a .env file in the explainroo folder:
OPENROUTER_API_KEY=sk-or-...
Git ignores .env. An environment variable with the same name works too. If you ask your agent for a video with images and there is no key, the agent asks you for one.
How is an image made?
node bin/explainroo.js image videos/jump-start cables "Two cars parked nose to nose with their hoods open, jumper cables between the batteries"
node bin/explainroo.js images videos/jump-start
The first command saves assets/cables.png. The second lists all the images made for this project and what each one cost. explainroo also writes each request with its cost to build/images/usage.jsonl.
| Option | What it does |
|---|---|
--model best |
OpenAI GPT Image 2, the default. It does what you ask. |
--model cheap |
Google Gemini 3.1 Flash Image. It costs about half, but often leaves out parts of the prompt. |
--aspect 16:9 |
the shape of the image: 16:9, 9:16, 1:1, 4:3, 3:4, 3:2 or 2:3. It follows the video by default. |
--ref assets/a.png |
sends an earlier image along, so a person or an object keeps the same look. Separate several files with commas. |
--no-style |
leaves out the default style |
Each image gets a default style. It is a flat illustration with soft colors, a plain light background and no text. You can set your own style for the whole video in video.json:
"images": { "model": "best", "style": "Watercolor illustration, warm colors, no text" }
Why does every image need a check?
Image models get details wrong. A hand gets six fingers, a cable goes on the wrong terminal, or a sign shows letters that look like words but are not. AGENTS.md tells the agent to open each image after making it. When something is wrong, the agent changes the prompt and makes the image again.
Two things help. Keep text out of images and put words on screen with s.text instead. And use one style for the whole video, so the pictures look like they belong together.