Build a consistent AI video in Venice Studio with reference-to-video
This guide walks through the Venice Studio end to end, using the Kling 3 reference-to-video workflow to produce a short, multi-shot commercial with consistent characters, scenes, music, and titles. You will learn how each tab fits together and the exact prompting approach that keeps your generations visually stable.
- What each of the five Studio tabs does and when to use it
- How reference-to-video keeps characters, objects, and scenes consistent across shots
- How to set up elements, identity locks, scene references, and frame control
- A repeatable prompt formula for video generation models
- How to assemble shots, add AI music, and export a finished video in the movie editor
Touring the five tabs of Venice Studio
The Venice Studio brings image and video creation into one workspace. Across the top you will find five tabs: image generation, image editing, video generation, video editing (the movie editor), and your assets library.
The image tab works the way most image tools do: you write a prompt, set resolution and style, pick a model, and the cost of the generation is shown before you run it. The image editing tab lets you upload a file or pull one from your library, then edit with a prompt, combine multiple images, upscale, or remove the background. Everything you generate in your browser through Venice collects automatically in the library tab, so any asset is available to reuse later.
The video and movie editor tabs are where this guide spends most of its time. Treat the library as the connective tissue: images you make become reference inputs for video, and finished clips become material for the timeline.
The video generation interface and Kling 3
In the video tab, start with model selection on the left. You can sort by name or speed and filter by category, including uncensored models and audio-enabled models. The available models change depending on whether you start from a text prompt, an image, or a video. For example, a model offered for image-to-video may not appear for video-to-video, so pick your starting input first.
Newer models like Kling 3 add the features that make consistency possible. You can supply a first-frame image and a last-frame image, attach element and scene reference images, and prompt multiple shots in a single generation. For each shot you describe what should happen and set its length, then generate. Kling 3 supports multi-shot output and caps a generation at 15 seconds total, so plan shot lengths to fit within that limit.
Reference-to-video: elements, scenes, and frame control
Reference-to-video means locking in the appearance of characters, objects, and scenes so your clips stay visually consistent, something that is hard to achieve otherwise. There are three kinds of visual input to understand.
Elements are the characters or objects you want to keep stable. You can use up to four per generation. Scene references define the stage: where the elements are set. Frame control sets where a clip starts and, optionally, where it ends, by uploading exact first and last frame images at the top of the generation form.
Each element should also get an identity lock: a frontal image plus up to three additional angles, so the model knows what the subject looks like as it turns and moves. The demo project uses three elements (a scholar, a merchant, and a compass) set across several scenes, and tags them consistently as "element one," "element two," and "image one" inside the prompts.
The mega prompt template for planning a project
Rather than write every image and video prompt by hand, the guide uses a free mega prompt template (linked in the video description) that turns a frontier chat model into a creative director for your project. Paste the full template into the Venice chat window. A stronger model such as Opus gives better results than a small open-source model, though either will work.
To use it, replace every placeholder: what the project is about, who the characters are, the objects, and the settings. Describe everything in detail and send it off. The model returns all the prompts you need, both for generating your element and scene reference images and for generating each video shot.
You then take the image prompts into the image studio first. For each character you generate a frontal plus three angle references; for each object you generate its angles; and you generate the scene settings. With those reference images in hand, you move to the video tab to assemble shots.
Generating shots and the prompt formula pros use
For each shot, attach the elements that appear in it along with their reference images, add the relevant scene image, and write the shot prompt using your element and image tags. If a character is not present in a shot, you can reassign which object occupies an element slot. When a scene reference and character are supplied, you usually do not need a first-frame image; the model generates the opening frame on its own. A convenient detail: you can queue a new generation while a previous one is still running.
A reliable shot prompt follows a clear structure:
[subject / element] + [action] + [environment from scene reference image] + [camera movement] + [lighting]
These models respond well to plain camera language like "slow camera push forward," "tracking shot from behind," "close up," "wide cinematic shot," or "camera hold static." Avoid complex moves that can overwhelm the model. Keep prompts brief but not vague: too short is ambiguous, too long invites contradictions. Place a single camera instruction early, keep vocabulary consistent, and always state spatial position (foreground left, entering from right, coming from the background). Toggle the generate audio button to add sound effects, and include sound cues directly in the prompt.
Editing the timeline and adding AI music
With your shots done, open the movie editor and use the media tab on the left to pull every generated clip onto the timeline in order. From here it behaves like standard non-linear editing software. You can add transitions such as a crossfade between shots, trim clips that run long, and cut out unwanted endings. If a transition freezes at the start of a clip, shortening that clip helps the crossfade land cleanly.
To unify shots that have different ambient audio, add a single background music track. In the base Venice chat, click the plus and choose generate audio. The guide uses ElevenLabs music with a one-minute length (the shortest available), pastes the music prompt from the mega prompt template, and sets it to instrumental only. Find the finished track in your media under the audio tab, drop it on the timeline, and mix the level so it sits under the dialogue. You can cut and rearrange the music, then add fade-ins and fade-outs to shape the final beats.
Adding a title, exporting, and what's next
Finish with a text title. In the demo, the imaginary product is a compass called Libera, so the title carries the product name and the slogan "Your direction, yours alone," positioned on screen where it reads well. When the edit is ready, click export. Venice renders all the clips together with the audio and title and produces a downloadable file.
The Studio is in an early stage, and the video notes you can also export your assets into desktop editors like Final Cut or iMovie for finer control. Updates are expected weekly with new functionality. The closing invitation is to grab the mega prompt for reference-to-video projects and share your work in the Venice Discord.
Key takeaways
- Venice Studio handles the full pipeline in five tabs: image generation, image editing, video generation, the movie editor, and a shared asset library.
- Reference-to-video keeps content consistent through three inputs: elements (with identity-lock angles), scene references, and first/last frame control.
- The mega prompt template plans an entire project, generating both image and video prompts so you build reference images before generating shots.
- A strong video prompt combines subject, action, environment, one early camera instruction, and lighting, with consistent tags and stated spatial positions.
- 214votes
- 97votes
- 142votes
Adapted from the @askvenice video on YouTube. Models and prices change fast; verify current details in Venice before production use.