How to create a consistent AI avatar in Venice with Seedance 2.5 and Wan 3.0
Post as a character from another world instead of as yourself, and keep that character's face and voice identical in every video. This guide builds a fictional knight living in Venice, with you as the actor, using Seedance 2.5 to reskin your look and Wan 3.0 for reliable lip sync.
- How to design a character and a short world bible in the agentic chat, and lock one image as your face reference
- Why the one-prompt shortcut drifts, and why filming yourself as the actor is the consistent route
- How to reskin your look with Seedance 2.5 in Venice Studio: face, voice, and background swaps
- How to lock a permanent voice with the voice changer, using a preset or an ElevenLabs voice ID
- How to fix a failed reference swap step by step instead of burning credits on re-rolls
- How to use Wan 3.0's built-in lip sync, and a cheap single-image talking-head fallback
Design the avatar and its world first
A consistent avatar starts with one locked character, so build that before you generate any video. Open the agentic chat and describe the persona in plain language. As a running example, we build a knight from a Renaissance world, living in Venice after being wrongfully excommunicated from his home country, who once slayed a dragon. Ask for three photo-realistic concepts, and because this is for social, set a 9:16 vertical aspect ratio up front.
Pick the concept that fits and save that image, because it becomes the face reference for everything you make later. While you are here, ask the same chat to write a short story and world bible for that concept, plus a social media strategy for the photos and videos the character will post. Go as shallow or as deep as you like on the backstory. What matters is that you leave this step with one character image and a world to place it in. This is how you show up online as a persona rather than as yourself.
The quick way, and why it doesn't hold up
The fastest route is a single instruction to the agentic chat: generate a 15 second video of the character introducing himself for Instagram. You get a lifelike, selfie-style clip in one shot, and it looks convincing on its own.
Two problems make it a dead end for an ongoing account. It often ignores your saved character image, so the face drifts from clip to clip, and the voice is invented by the model, so it will not match next time. That is fine for a one-off test to see how the idea feels, but consistency of face and voice is the whole point of an avatar, and this method gives you neither over the long run.
The reliable way: film yourself, then reskin with Seedance 2.5
The dependable approach is to be the actor yourself. Record a normal selfie video of you delivering the lines, then reskin your look into the character. In Venice Studio, open the Video tab and drag your recording in as a video reference, then add your locked character image as an image reference.
Use the edit-video option and say exactly what to change: make the man on screen look like image one, give him a deeper, gravelly, mysterious voice, and change the background to a Venetian palazzo with a canal behind him. Keep your aspect ratio at 9:16 and match the clip length, here 15 seconds. This edit runs on Seedance 2.5, and the result is strong: your performance, the character's face, a new setting.
The catch is the voice. Seedance invents a voice for the character, and you cannot reproduce that exact voice on the next generation. You might get lucky for a few clips, but across an account the character will slowly drift out of sound even when the face holds. That is the problem the next two steps solve.
Lock a permanent voice with the voice changer
To keep the voice identical every time, stop letting the video model invent it. Use the voice changer in Venice Studio instead. Upload the audio from your selfie recording and choose a target voice: either one of the presets, or a custom voice by pasting an ElevenLabs voice ID. To get a custom ID, open the ElevenLabs voice library, open a voice's menu, copy its voice ID, and paste it into the voice changer. This example uses a preset called Daniel.
Run it, download the result, and save it to your assets library. You now have one reference voice file for the character. Feed that same file into every future generation and he sounds the same in every video, with no luck involved. This single audio file is what makes the persona repeatable.
Prompt all three references, and fix a failed swap step by step
Now you can drive a generation from three references at once: your video, the character image, and the reference voice file. Reference each by number in the prompt. We use this one:
Transform the man in video 1 into the exact likeness of image 1, matching his
facial structure, features, and presence seamlessly.
Next, change his voice into the exact voice of audio 1, matching the tone, rhythm,
enunciation, and presence of the voice seamlessly.
Finally, transport the scene to the golden glow of a Venetian palazzo at dusk,
where weathered limestone walls reflect on the dark, shimmering canal waters behind him.
An image reference of the setting helps too, since a picture carries more than a sentence of description. Generate one in the agentic chat, add it as image two, then set your duration, aspect ratio, and resolution and generate.
Here is the honest part: sometimes a reference just does not take. The face and background swap land, but the voice comes back as your own, or as a completely new generation. When that happens, do not keep re-rolling the whole prompt and burning credits on what is close to a coin flip. Go step by step. Fix the voice on its own in the audio studio: run your recorded audio through the voice changer with the same target voice, then attach that corrected file. It is an extra step, but it is cheaper and more reliable than hoping one large prompt gets everything right at once.
If you would rather keep prompting, the other route is to take the prompt back to the agentic chat, say precisely what failed, for example that the audio reference swap is not working, and let it rewrite the prompt. Switching models there gives you different rewrites, often with a sharper instruction like "regenerate the man's dialogue audio from scratch, matching audio 1's voice." It can get you there eventually, but it costs more and it is not guaranteed.
Reliable lip sync with Wan 3.0, and a cheap talking-head fallback
When the voice swap keeps fighting you, switch models. Wan 3.0's reference-to-video has lip sync built in, which makes it the more dependable tool for matching audio to a face. Test it three ways from the same material and compare the results: a full face-and-voice swap, a voice-only swap, and a voice swap against a single image instead of a video. Here, Wan 3.0 syncs the reference audio to the original clip cleanly.
The cheapest version needs only two references. Record a line, run it through the voice changer into your character's voice, then sync that audio to one still image. A simple prompt does the job: sync audio one to image one, where image one is the character in a scene roughly the length of the audio. Cleaned up in the agentic chat, that instruction becomes something like "animate image one as a talking head, lip sync driven by the person in image one speaking audio one's dialogue with accurate mouth shapes." When budget is tight, this image-plus-audio talking head is the fallback that still stays on character.
Keep it consistent: reuse one prompt template
The habit that protects everything above is to settle on one prompt template and reuse it unchanged. Reordering a few words can swing the result, so once a prompt reliably gives you the face, voice, and setting you want, save it and run the same structure every time. That is what turns a lucky good clip into a character you can post as indefinitely.
One note the creator makes plainly, and it is worth repeating: use this to create, not to deceive. Playing a fictional character is fun, but using an avatar to spread misinformation or to impersonate a real person can land you in real trouble. Build the persona, and be mindful of what you do with it.
Key takeaways
- Lock one character image and a short world bible in the agentic chat before you generate any video.
- Be your own actor: film yourself, then reskin your look with Seedance 2.5 in Venice Studio's Video tab.
- Make a permanent reference voice file with the voice changer, using a preset or an ElevenLabs voice ID, and reuse it every time.
- When a swap fails, fix it step by step in the audio studio instead of re-rolling a whole prompt and burning credits.
- Use Wan 3.0 for reliable lip sync, and fall back to a single image plus an audio file when budget is tight.
- Save one prompt template and reuse it unchanged so the face and voice stay identical across every post.
Adapted from the @askvenice video on YouTube. Models and prices change fast; verify current details in Venice before production use.