A small, auditable workflow for comparing AI video prompts
When two generated clips look different, it is tempting to credit the new prompt. But a changed model, input image, duration, or random sample can explain the difference too. A useful first step is to make the comparison traceable.
This tutorial describes an experiment design, not a performance result. The accompanying examples are original, AI-assisted synthetic prompts. They have not been run or scored, and no model-quality claims follow from them.
Start with one pair
Use the same base scene in both variants:
Use an owned reference photo of one plain ceramic mug on a wooden table. The mug stays still, with its shape and surface pattern unchanged. Soft window light remains constant.
Append one of these camera instructions:
A: Locked camera, medium close-up.
B: Slow camera push-in from a medium close-up.
The change is deliberately small. The task is to inspect camera behavior while watching for unintended changes to the mug. This does not guarantee the model will preserve the scene or obey either instruction.
Record the run before judging the clip
A minimal record should identify the prompt, provider, model and exposed version, run date, seed if available, duration, aspect ratio, resolution, input reference, output, and generation success. Leave unavailable settings blank rather than guessing them. A hash can identify a reference file without publishing the image itself.
For a pair, keep the reference image and available settings fixed. Matching seeds, where supported, improves traceability but does not promise reproducibility. Repeat independent samples when the budget permits; one impressive output is not evidence of a general advantage.
Separate prompt adherence from visual stability
Inspect the requested movement first. Did the camera actually push in? Then inspect unintended changes: did the handle split, the mug change shape, or the scene jump? These are different observations and deserve separate fields.
A simple manual rubric can use three labels: 0 for absent or contradicted behavior, 1 for partial behavior, and 2 for clear behavior. For temporal consistency, use 0 for severe discontinuity, 1 for some visible instability, and 2 for no obvious instability in the reviewed clip. These are proposed ordinal labels, not a validated objective metric.
Do not assign a perceptual score to a failed generation. Keep the failed run in the record and describe what happened. If several people review outputs, randomize their order and keep disagreement notes.
Download the complete study pack: CSV, JSONL, empty observations sheet, and documentation.
Keep a compact data format
The study pack uses one CSV row per prompt and a separate observations CSV for results. Its 24 prompts form 12 pairs across camera movement, motion, lighting, focus, and framing. Half of the pairs use text-to-video; half require an owned or licensed reference image.
All included records have generation_status=not_run. The observations file contains headings only. Separating the design from the results helps prevent an untested example from being mistaken for an evaluated sample.
Publish the limitations with the examples
This is a small English-only convenience sample of simple scenes. It includes no human dialogue, no real-person identities, no generated outputs, and no reference images. A change in wording can affect more than the intended visual variable. Provider settings and model behavior also vary.
If you later share results, publish the sample count, failed runs, exposed settings, rating procedure, and rights to any shared media. Do not rank models using this template alone.
Disclosure: this resource was prepared with AI assistance for the Seadanse project, an independent web interface for AI video generation. The workflow can be used with any compatible service; Seadanse is not required. The original study-pack text and data are offered under CC0 1.0; third-party inputs and outputs have their own rights and terms.
