Use text-to-video when an explainer needs moving images that set a mood or illustrate an idea. Use code-rendered video when viewers need to read exact labels, follow diagrams, or understand a product interface. Text-to-video models turn written prompts into moving images. Code-rendered videos use elements such as text, shapes, charts, and screenshots.
explainroo is an open-source kit for making code-rendered explainers and product demos. An AI coding agent uses it to create and check a narrated video on your computer.
Text-to-video vs. code-rendered video
A text-to-video model creates footage from a prompt, such as “a small business owner reviewing a shipment in a bright warehouse.” It turns the prompt into a moving scene. This works well when mood or setting matters more than exact details on screen.
A code-rendered video defines what appears in each scene and when. As the narration reaches each point, a scene can reveal labels; it can also show a chart or highlight a line of code. The renderer combines the scenes, voice, and other audio into a video file. Explicit scene instructions let the creator control wording, placement, and timing.
The choice matters most when viewers need to learn from what they see. A prompt can suggest a “rising sales chart,” while a code-rendered chart can show the figures and labels described in the narration.
When to use text-to-video or code-rendered video
Use text-to-video for illustrative scenes, atmosphere, and visual variety. Use code rendering for step-by-step explanations, diagrams, charts, and interface walkthroughs. Choose the method based on what viewers need to understand, not just on which visuals look new.
Use generated scenes for atmosphere
Generated footage can establish a setting or illustrate an abstract idea. For example, an explainer about supply chains could open with a generated warehouse scene, then show a diagram of how goods move between locations.
The tradeoff is precision. Prompts cannot reliably control exact wording, object placement, or continuity between shots. A generated scene can communicate “busy warehouse” without showing the specific package, sign, or action named in the script. Check each clip for accuracy and continuity before using it.
Use code-rendered scenes for labels and steps
Use code rendering when viewers need to read exact text or follow a sequence. A product explainer can highlight a button, show someone completing a form, and reveal the result. A science lesson can add labels to a diagram as the narrator defines each part.
Someone, or an AI coding agent, must define the scenes and visual rules. The video can feel rigid if the creator does not plan the layout, pacing, and visual hierarchy. Start by showing only the information that the narration explains in each scene.
Control, consistency, and revisions
Code-rendered scenes preserve details more reliably because instructions define the text, layout, and timing. The creator can change a wrong label or adjust a cue if a diagram appears too early. These edits may require technical work, but they target a specific part of the video.
Changing a prompt can alter more than the intended detail. A prompt revised to fix an object may also change the scene’s composition or appearance. As a result, characters, settings, and objects are harder to keep consistent across clips.
With either method, keep the script and visuals in sync. When the narration changes, check that the scene timing still fits. If the explanation depends on exact figures or product wording, check those details against the source before rendering or approving the footage.
Production effort and review
Review generated footage for continuity and fit with the narration, including unintended details. For code-rendered scenes, assess accuracy and timing, and make sure the text is readable. Watch the final video at its intended size. Labels that look clear on a large monitor can be hard to read on a phone.
Code rendering makes specific revisions more predictable, but designing scenes takes planning. Text-to-video creates varied images from prompts, but revising prompts and reviewing clips takes time. For either method, compare first-draft time with the effort required to produce a clear, accurate video that is ready to use.
Make a code-rendered explainer with explainroo
explainroo is a free, open-source kit that works with coding agents such as Claude Code, Codex, Pi, OpenCode, and Gemini CLI. It works best with a compatible coding agent. To start, give the agent your topic and ask it to use explainroo, clone the GitHub repository, read AGENTS.md, and follow the instructions.
The agent writes two files: script.md for the narration and scenes.js for the visuals. A scene can appear when a chosen word is spoken. Kokoro, an open voice model, reads the English script in an American or British voice. Whisper identifies when each word is spoken, so scenes can change with the narration.
Chrome draws each frame on an HTML canvas. Scenes can include charts, code, screenshots, and Lucide icons. Rough.js adds hand-drawn-looking lines. explainroo also creates background music and sound effects, lowers the music while the voice speaks, and uses FFmpeg to assemble the MP4.
Before finishing, the agent checks still images of the scenes and a contact sheet. It also checks text for cutoffs or overlaps, along with speech for misread words. The agent cannot watch the video, so review the MP4 yourself too. In video.json, you can choose from paper, clean, chalk, blueprint, and midnight looks. Formats include wide, tall, and square layouts. Some formats offer word-by-word captions. The pace setting in video.json adjusts the narration, pauses, and animations.
For a product demo, tell the agent which product to show and where to find its code or website. The agent rebuilds the screens with the product’s colors, fonts, and button labels. It then shows a pointer clicking and typing as the narration explains the interface.
The kit cannot edit existing camera footage or create talking avatars or AI-generated live-action footage. It also has no drag-and-drop editor. Optional AI illustrations use OpenRouter and cost money per image; the rest of explainroo is free per video. The coding agent is a separate service with its own pricing. To run explainroo locally, you need Node.js, FFmpeg, and Chrome or Chromium. Initial setup downloads the voice and timing models. You do not need a graphics card.
Choose by the information shown
Start with what viewers need to understand on screen. If the explanation depends on an exact label, number, diagram, or interface, render it from code. If it depends on a setting, mood, or illustrative action, use generated footage for that part.
A hybrid approach gives each method a clear role. Use generated clips to set the mood, then use code-rendered titles, diagrams, or data for facts viewers must read. Do not put important claims in footage that cannot reliably preserve exact wording. For a narrated explainer that needs accurate visuals, explainroo offers a code-rendered option. Use a separate text-to-video service when you need generated moving images.
Key Takeaways:
- Choose by the information viewers must learn: Use code rendering for exact labels, charts, steps, and interface details; use text-to-video for illustrative scenes and atmosphere.
- Review the final video: Check generated footage for continuity and relevance. Check code-rendered scenes for accuracy, readability, and timing.
- Combine methods when their jobs differ: Use generated clips to set a scene and code-rendered visuals to explain facts precisely.
- Use explainroo for narrated, code-rendered explainers: An AI coding agent creates the script and scenes. explainroo renders and checks the video on your computer.
Conclusion
Text-to-video offers visual variety. Code rendering gives you direct control over what viewers see and when. Choose based on whether the explainer needs to evoke a scene or communicate precise information.
For narrated explainers and product demos that need exact visuals, explainroo uses a coding agent’s script and scene definitions to make an MP4 and run automated checks. Review the finished video yourself. If atmosphere matters too, combine code-rendered explanations with generated clips that do not contain critical text or data.