Explain a Paper home

Built, not generated

How we keep a video true to the paper.

AI video models paint every frame from patterns they have learned, which is why they can show a number, a label or a chart that is not in the paper. We use AI to help write the script, and then our software draws every frame from that script and the paper's own data.

How we keep videos accurate (1:20, captioned). About this video and its transcript

What an AI video model does

A text-to-video model is trained on a very large collection of videos and their descriptions. From them it learns what things tend to look like, and how they tend to move. Given a prompt, it does not look anything up. It starts from random noise and, over many small steps, turns that noise into frames that fit the patterns it learned. This is called diffusion, and most current AI video models combine it with a transformer, the same kind of network behind language models. OpenAI's technical report on Sora describes exactly this: a diffusion model that, "given input noisy patches (and conditioning information like text prompts)", is "trained to predict the original 'clean' patches". Google's model card for Veo 3 calls latent diffusion "the de facto standard approach for modern image, audio, and video generative models".

So every pixel is a prediction. The model is painting what a frame like that usually looks like, not drawing from a source. For footage of the real world that works remarkably well. For a research video, it is the wrong tool, for four reasons.

Why text, numbers and charts come out wrong

  • Text. To the model, letters are shapes. It learns what writing looks like, not how to spell, so on-screen words often come out garbled or invented. A 2025 benchmark that tested ten text-to-video models on on-screen text, from captions to formulas, found that "most struggle to generate legible, consistent text".
  • Numbers. Nothing in the way a "7" looks tells the model whether 7 is the right number. A benchmark of compositional prompts found that models do well with fewer than three objects but "often fail to accurately generate larger quantities".
  • Charts. The model has never seen your data. Asked for a chart, it paints something chart-shaped: axes, a line, some points, a heading. Nothing ties those points to the paper's table.
  • Physics and continuity. Each frame is predicted, so details can drift from one frame to the next, and objects can appear, vanish or move impossibly. OpenAI said of Sora that it "does not accurately model the physics of many basic interactions, like glass shattering". Google's Veo 3 model card says that "maintaining complete consistency throughout complex scenes or those with complex motion, remains a challenge". Independent benchmarks of physical common sense reach the same conclusion, and one study found that AI video models copy the closest example they were trained on rather than learning the underlying law.

When a model produces something that looks right but is not true, it is often called a hallucination. In a research video that is the one thing that cannot be allowed through, because a viewer has no way to tell a painted number from a real one.

The makers say so themselves. The team behind Wan, a leading open AI video model, write that its "performance in specific localized scenarios, such as education and medicine, may be insufficient". A 2025 benchmark for turning research papers into videos found that a leading end-to-end AI video model gave blurred on-screen text and left out much of the paper's content.

What we found when we tried it

In September 2026 we ran an open AI video model on our own computer, with prompts taken scene by scene from our video of Hubble's 1929 paper. Asked for the paper's title page, it drew a page of pseudo-text and spelled the author "Eovin Hubble". Asked for a chart of Hubble's 24 nebulae, it drew one line, two dots and a garbled heading, with none of his data. Asked for our web address, it wrote "expllinapaper.com". A few short headings came out right, but no page of text, no table and no chart did.

What we do instead

Our videos are AI-facilitated, not AI-generated. Our specialised AI agents help with the writing; everything you see is drawn by our own software, the way digital animation has long been made.

Two video frames showing the same chart. Left, painted by an AI video model: the title and axis labels are letter-shaped scribbles and the points are smeared blobs, because such a model can invent a number, a label or a result. Right, built from the paper: the typeset title Hubble, 1929: 24 nebulae, axes labelled Distance and Velocity, and 24 sharp points plotted from the paper's own numbers.
An illustration of the difference. The built chart is from our pilot video: the 24 nebulae of Hubble's 1929 Table 1.

1. A script you approve

Our specialised AI agents draft the script from your paper and send it to you. Each line notes the page, table or figure it comes from. Then it is yours to edit, and nothing is drawn until you approve it. You can add human review: one of our editors checks it against the paper.

2. Every frame drawn

Our software draws each frame in code. Words are typeset, charts are plotted from the paper's own numbers, and diagrams are drawn to show what the script says. No AI image or video model is involved.

3. A synthetic narrator

You choose from 28 high-quality synthetic voices in English, British and American, women and men, and the one you choose reads the script you approved. The voices run efficiently on our own machines, and the narrator is disclosed as synthetic in the first caption, the transcript and the video file.

What that means for accuracy

  • No made-up facts: our software draws only what is in the script you approve, and the script notes the page, table or figure behind each line, so you can check it against the paper.
  • Every word on screen is typeset from the approved script, so there is no garbled lettering.
  • Every chart is plotted from the paper's own numbers. In our pilot video of Hubble's 1929 paper, all 24 nebulae come from his Table 1.
  • The same scene looks the same each time it is drawn, so nothing drifts from frame to frame.

That is a description of the method, not a promise that a video cannot contain a mistake. Our specialised AI agents draft the script, and they can misread a paper. That is why each line notes its source, why you approve the script before anything is drawn, and why human review is there if you want a second person to check it. We do not run automated accuracy checks, and we do not claim to. What the method does is make sure the pictures add nothing the script does not say, so checking the script is checking the video.

It is also lighter on electricity. In our September 2026 measurement, a whole 70-second video, the AI script drafting included, used about 50 to 90 watt-hours. One take of AI-generated footage for the same video, on the same machine, used 1.7 kilowatt-hours.

Where AI video models shine

AI video models are a real achievement. They can produce photo-real footage of people, places and weather, camera moves that would be costly to film, and mood pieces in many visual styles, and they are improving quickly: their makers report better physics and more faithful prompt following with each release. For a short advert, a storyboard or a shot of a stormy coastline, they may well be the right tool.

A research explainer needs something else. Its job is to get the paper's words, numbers and findings across exactly, and to let anyone check them against the source. Even a frame that happens to come out right is a painting of a result, with nothing linking it back to the paper. A frame we draw is the result, plotted from the paper and traceable to the line of the script that called for it.

More on how we use AI: AI and ethics. What our engine has made, and the electricity we estimate that saved: our numbers.

Sources

Sources checked on 1 October 2026. Quotations are the authors' own words. The models move quickly, so a newer release may do better than the figures quoted here.