One Speaker, Three Languages, One Lip-Sync Test
A spare multilingual lip-sync launch demo that closes with a matched side-by-side comparison.
Field notes / why it works
The capability claim is visible at 0:00 and the first transformed speech starts by 0:02.1. The fixed speaker, white backdrop and tiny language/source labels make the German, original English and French passages comparable at about 0:02, 0:15.6 and 0:30.2. A two-up baseline-versus-product test at 0:43.3 gives the 56-second demo a concrete payoff.
Plate 01 / The format in one breath
One person in a locked, bright studio composition demonstrates a visual AI capability through several language variants. Brief labels tell viewers what changed, while unchanged wardrobe, backdrop and framing make the result itself the evidence. A late side-by-side comparison turns the claim into a test.
Why it works
- The opening promise is visible at 0:00: a restrained headline beside the subject says the speaker's style will survive lip syncing. The first example begins at 0:02.1.
- Around 0:15.6 the label changes from a generated language version to the original English clip; around 0:30.2 it changes to another language. This gives viewers repeated A/B checks within the same scene.
- At 0:43.3, two face crops appear side by side and are labeled as a competitor category and this product. The contrast is the payoff, rather than an unsupported superlative.
Plate 02 / Format card
| Platform · aspect · length | X · 16:9 · 1280×720 · 29.97 fps · 56.1 s |
| Pace | 4 measured shots, 3 cuts, 3.2 cuts/min; 16 in-shot changes, 20.3 total picture changes/min; mean shot 14.03 s, median 6.69 s |
| Script | 134 machine-transcribed words, 153 wpm overall and 185 wpm while speaking; mixed languages make exact wording unreliable |
| Voice | Single low-register adult speaker archetype, measured median 100 Hz; moderately varied delivery |
| On screen | Nearly all time is one live-action-looking seated speaker; the final ~11 s uses a two-up face comparison; ~1.5 s is a white legal card |
| Captions | No spoken-word captions; small, static lowercase proof labels and language tags |
| Sound | Speech leads. A tonal bed is measured, but whether it is separate music is unclear; no distinct cut effects evidenced. −15.4 LUFS integrated, −0.3 dBFS true peak. |
| Structure | Promise → language proof → original control → second language proof → two-up comparison → fine-print stop |
Pro / 13 plates + recipe
The first plates are open. The full dissection is Pro.
Get the timed shots, type and sound specs, and the runnable recipe.
Open the full blueprint- Plate 03 / The hook, frame by frame (first 3 seconds)Pro
- Plate 04 / Structure (beat sheet)Pro
- Plate 05 / Shot grammarPro
- Plate 06 / Visual style guidePro
- Plate 07 / Voice and deliveryPro
- Plate 08 / Sound designPro
- Plate 09 / PackagingPro
- Plate 10 / Script templatePro
- Plate 11 / Production planPro
- Plate 12 / PromptsPro
- Plate 13 / Edit recipePro
- Plate 14 / QA checklistPro
- Plate 15 / Make it yoursPro
- Recipe / runnable promptPro