Comparing nine AI video models on the same prompts to build a cyberpunk scene
This guide walks through a hands-on experiment that ran nine identical shot prompts through nine video models inside Venice, then assembled the best results into a film noir teaser. You'll learn how to evaluate AI video output, write prompts that hold a consistent style, and pick the right model shot by shot.
- A four-part framework for rating AI video generations
- How generation speed compares across the models available in Venice
- How to write detailed, style-driven prompts that survive across different models
- Why testing multiple models per shot beats reprompting a single model
- How to assemble shots, fix color inconsistencies, and add narration
The experiment and why use multiple models
The project: 81 videos generated from nine prompts run through nine different models, costing over $100, all to assemble a single cyberpunk movie teaser. Each prompt described one shot of a film noir detective story, and the goal was to find the single best version of each shot across every model.
Venice puts all nine models in one interface, which makes this kind of side-by-side iteration practical. Instead of committing to one tool, you can run the same prompt everywhere and compare. The throughline of the whole experiment is simple: different models interpret the same prompt differently, so having them in one place lets you fish for the best result rather than fighting one model over and over.
Generation speed across models
Speed was averaged across all nine prompts for each model. The fastest was Wan 2.2, followed by Veo 3.1 fast and Veo 3 fast. The slowest were Sora 2 Pro, Wan 2.5 preview, and Sora 2.
Speed alone does not decide quality. Veo 3 fast turned out to produce some of the best results in the experiment while also being among the quickest, so it offers a strong balance. Treat these numbers as one data point: your results will vary by prompt and load.
The four-criteria evaluation framework
Rate every generation from 1 to 5 on four axes:
- Prompt accuracy: Does the video show what you asked for? Check shot composition, lighting, camera angle, movement, character positioning, atmosphere, and specific details like rain streaks or neon.
- Visual consistency: Is it coherent throughout, with no morphing objects, physics violations, elements popping in or out, warping textures, or temporal glitches?
- Aesthetic quality: Does it look cinematic? Judge lighting, mood, color grading, contrast, framing, depth, and overall appeal.
- Technical quality: Look at resolution, sharpness, motion blur, frame-rate smoothness, compression artifacts, grain, and edge definition.
Two best practices make scoring honest. Watch each clip at least twice, once for a gut impression and once for detail. And grade relative to prompt difficulty: simpler prompts deserve higher standards because the model has fewer constraints, while complex prompts are harder to satisfy. Compare same-scene clips back to back.
The nine models at a glance
The lineup spans open-source and frontier models, each with different strengths.
Wan (2.2 and 2.5 preview) are open-source models that handle audio-visual sync with character dialogue and produce rich video dynamics. Kling 2.5 Turbo Pro is fast and well priced, has no audio, but excels at character animation and movement. Google Veo 3 and Veo 3.1 are both strong, and the experiment repeatedly shows the older Veo 3 outperforming the newer 3.1 on certain shots. Sora 2 and Sora 2 Pro bring ultra-realistic physics and generate whole scenes with audio rather than a single isolated shot.
The takeaway from the introductions is that newer does not always mean better for a given shot, which is exactly why testing across models pays off.
Running the shots: prompting and picking winners
Each shot was a detailed prompt specifying composition, lens, lighting, mood, and concrete details, then compared across all nine models. A few patterns emerged.
The cigarette smoke test (a static composition) exposed physics errors, with several models burning the cigarette from both ends or placing it backwards; Veo 3 fast and Wan 2.2 produced the cleanest, most realistic results. The detective-at-desk shot used style stacking, naming three films in the prompt to fuse their aesthetics:
Wide shot through a rain-covered window of a lone detective sitting at a desk in a dimly lit office. Aesthetic combining Se7en oppressive atmosphere, Heat urban isolation, Blade Runner 2049 neon cityscape. Foreground raindrops in soft focus, middleground detective backlit by a single desk lamp, background a blurred noir cityscape with cool blue and warm orange light. Static composition with vertical blinds casting diagonal shadows. Melancholic isolated mood with strong color contrast between warm interior and cool exterior.
A key prompting lesson: the simpler "detective walking in the rain" prompt should be judged more harshly because it gives the model freedom, while heavily constrained prompts (the chase shot specifying the woman "20 to 30 feet ahead," tense cat-and-mouse atmosphere) test whether the model can follow exact direction. Sora 2 nailed the chase composition; Veo struggled and had characters walking toward each other instead of in pursuit.
Two more guides surfaced. Prompts carry no memory of each other, so a phrase like "from the detective's point of view" fails when that prompt never established a detective. And not every prompt will generate: the pistol-loading shot returned a censored response on Veo 3 full quality because of the gun, while other models produced floating weapons or triggers firing without a finger. When that happens, switch models.
Assembly, color matching, and narration
Once the winners were chosen, the shots went into a video editor. For newcomers, DaVinci Resolve has a capable free, multi-platform version; iMovie works on Mac and ClipChamp on Windows.
Playing the shots in sequence revealed problems no single clip showed in isolation. The teaser leaned on a blue-and-orange palette, and a couple of winning shots were missing the blue or skewed too sepia. That pushed a swap back to other model versions purely for color consistency, including switching to Veo 3 full quality and Sora 2 on specific shots. A hat that appeared on the detective in one clip broke character consistency, so that shot was replaced with a Sora 2 Pro version without the hat, and the rotary phone shot was moved earlier to where it made narrative sense.
Final narration was generated with ElevenLabs and layered over the cut. The honest reality of the workflow: even after heavy iteration the piece is a loose, mysterious teaser, not a tightly defined story. If you need precise continuity, reprompt with explicit details like "detective with no hat."
The multi-model strategy
The finished teaser holds together as a single aesthetic even though its shots came from open-source Wan, Google Veo, Sora, and others. That cohesion comes from detailed, style-heavy prompting, which buys you forgiveness when mixing models.
The closing pro tip: you do not have to run every model on every shot, but understand that each model produces a different interpretation of your prompt. When one model isn't landing the shot, trying a different model is often more productive than reprompting the same one repeatedly.
Key takeaways
- Score every generation 1 to 5 on prompt accuracy, visual consistency, aesthetic quality, and technical quality, watching each clip at least twice and grading against the prompt's difficulty.
- Detailed, style-driven prompts, including stacking multiple film references, give the most control and keep results consistent even across different models.
- Each prompt is stateless: it has no memory of other shots, so spell out characters, wardrobe, and point of view explicitly to maintain continuity.
- Different models interpret the same prompt differently, so testing across models beats reprompting one, and some shots only come together during assembly when color and character consistency become visible.
- 214votes
- 97votes
- 142votes
Adapted from the @askvenice video on YouTube. Models and prices change fast; verify current details in Venice before production use.