Voices from Prompts: Eleven v4 Launch Film
An audio-first voice-model launch film that proves expressive range through three performed scenes over restrained particle portraits.
Field notes / why it works
The first spoken word begins at 0:00 over near-blank white, and a small proof sentence becomes readable by about 0:02, turning attention toward the sound. The strongest replay peak at 0:12.6–0:14.1 coincides with the first particle portraits and performed scene; later examples at 1:07 and 1:44 show narrative range and practical dialogue. Only 3 clear cuts across 156 seconds leave room to hear changes in timing, emotion, and precision.
Plate 01 / The format in one breath
An almost empty screen makes a proof promise while a voice begins speaking. The film then lets several distinct, emotionally performed voices carry little scenes; tiny visible prompt directions and restrained particle portraits show what drives them. Brief abstract interludes name capabilities before another practical dialogue demo and a quiet brand close.
Why it works
- The proof claim starts with the first spoken word at 0:00, while the sentence gradually becomes legible by about 0:02. The sparse screen makes listening the task.
- The first replay peak, 0:12.6–0:14.1, lands exactly as a minimal audio mark gives way to the first two particle faces and performed dialogue.
- The film changes the kind of proof: theatrical emotional delivery around 0:15–0:50, then a narrative voice around 1:07–1:29, then a service call around 1:44–2:28. This demonstrates range without fast cutting.
Plate 02 / Format card
| Reference | |
|---|---|
| Platform · aspect · length | YouTube · 1280×720, 16:9, 30 fps · 156 s |
| Pace | 4 detected shots; 3 clear cuts, or 1.2–1.9 cuts/min counting two uncertain changes. Six internal visual steps; about 3.5 picture changes/min. Mean detected shot 39 s; one composition runs about 100 s. |
| Script | 279 transcribed words; 137 wpm over the film, 169 wpm while speaking; 38 short sentences averaging 7.3 words. Three questions. |
| Voice | Ensemble of synthetic performed voices: neutral guide, theatrical scene partners, animated broadcaster, calm service caller/agent. Medium measured median pitch (149 Hz) across the mix with unusually wide pitch movement. |
| On screen | Entirely designed graphics: mostly warm white field, tiny centered/lower text, dissolving point-cloud faces, brief saturated abstract colour, closing wordmark. No visible presenter or screen recording. |
| Captions | Small lower-center, two-line prompt/performance text; muted grey with selective salmon emphasis. These show performance directions and lines rather than large social subtitles. |
| Sound | Voices in the foreground; a continuous tonal, music-like bed measured about 3.5 dB below voice. Spoken pauses and performance effects carry the changes. Integrated −11.7 LUFS; measured true peak +1.1 dBFS. |
| Structure | Proof premise → expressive scene → capability turn → narrative voice scene → capability bridge → low-latency dialogue demo → brand close. No spoken CTA or loop. |
Pro / 13 plates + recipe
The first plates are open. The full dissection is Pro.
Get the timed shots, type and sound specs, and the runnable recipe.
Open the full blueprint- Plate 03 / The hook, frame by frame (first 3 seconds)Pro
- Plate 04 / Structure (beat sheet)Pro
- Plate 05 / Shot grammarPro
- Plate 06 / Visual style guidePro
- Plate 07 / Voice and deliveryPro
- Plate 08 / Sound designPro
- Plate 09 / PackagingPro
- Plate 10 / Script templatePro
- Plate 11 / Production planPro
- Plate 12 / PromptsPro
- Plate 13 / Edit recipePro
- Plate 14 / QA checklistPro
- Plate 15 / Make it yoursPro
- Recipe / runnable promptPro