Venice keysVeniceLearn · GuidesOpen Venice ↗
Guide · Video

Localize one video into seven languages with Happy Horse 1.1 on Venice

Take one set of reference photos and one short line, then generate the same speaker saying it with native, lip-synced audio across all seven languages Happy Horse 1.1 supports, each shot in front of a landmark from where that language is spoken. Your face, voice, and script stay on Venice the whole time.

What you'll learn
  • Why Happy Horse 1.1 owns multilingual talking-head video
  • How one reference image and a locked frame hold a speaker steady while the language and backdrop change
  • How to pick a single reference photo and a dead-center framing
  • How to generate the first clip, then localize it into all seven languages in front of a landmark each
  • How to draft cheaply on 720p and finalize keepers on 1080p

Why Happy Horse 1.1 for multilingual video

Most video models either skip audio or only lip-sync English. Happy Horse 1.1 generates picture and sound in one pass, so dialogue, room tone, and lip motion come out together instead of being stitched on afterward. Its edge is language: phoneme-level lip-sync in seven languages, with the lowest reported word error rate of any production video model.

The seven supported languages are English, Mandarin, Cantonese, Japanese, Korean, German, and French. Other languages still render, but the mouth match is only trained for these seven. Reach for Happy Horse when you want the best-looking SFW talking head or you need one message in many languages. For native multi-shot camera work, use Kling 3.0; for mature content or maximum freedom, use Seedance 2.0 or Wan 2.7.

The two levers, kept separate

One mistake causes most drifting faces and missed lip-sync: treating identity and dialogue as the same control. They are two levers, and they do two jobs.

One reference, locked framing, swap the line, the language, and the backdrop, regenerate. The speaker holds dead-center while the landmark behind them cuts from Paris to the Great Wall. That contrast is the proof: the same person, in the same spot, while the whole world changes. Two independent generations will not be pixel-identical, so prompt a static camera every time and nudge alignment in the editor if a clip drifts.

Pick one reference image and lock the framing

Choose one clean, front-facing, well-lit photo of the speaker, centered. Then set a medium-shot, dead-center, static-camera framing you will reuse for every clip. The model accepts up to nine references for holding a face across varied angles, but because the framing is locked here, one good photo is enough. The photo and the framing never change once you start.

One front-facing photo of [NAME]: even soft lighting, plain
background, calm friendly expression, head and shoulders centered.
Framing for every clip: medium shot, dead-center, static camera.

Generate the first language

In the Video tab, pick Happy Horse 1.1 (reference-to-video), attach your one reference photo, and set 720p, 16:9, audio on. Build the prompt from a skeleton you will reuse for every language. The language name, the quoted line, and the backdrop change between clips; your description of the speaker and the framing stay identical. Name the language before the quote, since that cues the lip-sync, and keep the line short with one speaker. Keep a dead-center, static framing so the speaker holds the same spot and the landmark reads softly behind them.

Medium shot of [NAME], centered in frame, static locked camera,
outdoors with [BACKDROP], soft daylight, shallow depth of field so the
landmark reads softly behind them. They say, in clear English: "[YOUR
LINE]" Natural blinking, subtle head movement, precise lip-sync.
Ambient: [AMBIENT], no music. Keep [NAME]'s exact face, hair, and
[WARDROBE] from the reference photo.

Paste a negative prompt to keep music and artifacts out, run Quote to see the cost, then Queue. Reshoot the line read once or twice on 720p and pick the best take.

background music, score, subtitles, captions, on-screen text,
warped mouth, frozen lips, out-of-sync audio, distorted face,
identity drift, extra fingers, oversaturation

Localize into all seven languages

Keep your one reference photo, your description of the speaker, and the dead-center framing identical. Change the language name, the translated line, and the backdrop, then regenerate at 720p. Here is one example line across all seven supported languages, with a suggested landmark for each. Swap in your own, and have a native speaker check each translation before you finalize, because awkward phrasing shows up on the lips.

English    (London, Tower Bridge):        "Privacy is a human right. On Venice, your ideas stay yours."
Mandarin   (the Great Wall of China):     "隐私是一项基本人权。在 Venice,你的创意永远属于你自己。"
Cantonese  (Hong Kong, Victoria Harbour): "私隱係一項基本人權。喺 Venice,你嘅創意永遠屬於你自己。"
Japanese   (Mount Fuji):                  "プライバシーは基本的人権です。Venice なら、あなたのアイデアはあなたのものです。"
Korean     (Seoul, Gyeongbokgung Palace): "프라이버시는 기본 인권입니다. Venice에서는 당신의 아이디어가 온전히 당신의 것입니다."
German     (Berlin, Brandenburg Gate):    "Privatsphäre ist ein Menschenrecht. Bei Venice gehören deine Ideen dir."
French     (Paris, the Eiffel Tower):     "La vie privée est un droit fondamental. Sur Venice, vos idées restent les vôtres."

Finalize and assemble

Pick the keeper for each language and re-render only those at 1080p by switching the resolution and regenerating. You only pay full resolution for the takes you keep, which is what makes the 720p draft pass worth it. Then open the Movie Editor, lay the clips in sequence, add a small sentence-case language label on each, and export at 1080p.

Key takeaways

Models and prices change fast; confirm the current Happy Horse model IDs, languages, and cost in Venice before production use. See the Happy Horse 1.1 model page for specs.