Eleven v3 voice performance reel
An audio-first product launch that demonstrates expressive synthetic voices over drifting gradients and visible scripts.
Field notes / why it works
The first dialogue begins at 0:08.6 and makes the product's output the evidence, while eight contrasting scenes reset attention at roughly 20–35 second intervals. The largest replay peak at 1:48.1–1:50.5 lands on the intimate self-aware voice passage, showing that expressive delivery, rather than visual spectacle, is the payoff. A continuous gradient canvas and 172 internal picture changes keep the 3:54 film alive without distracting from listening.
Plate 01 / The format in one breath
This is a product launch film that lets the product's output carry the pitch. An 8.6-second logo/reveal opens it; then a sequence of performed voice scenes demonstrates dialogue, emotion, character, and accent control. Small typeset scripts and slowly shifting luminous gradients keep the viewer oriented without competing with the sound.
Why it works
- The first voice at 0:08.6 starts a dialogue rather than a feature list; by 0:15, the voices are performing the capability being discussed.
- Chapters reset the listening task roughly every 20–35 seconds: sports at 0:54, character monologue at 1:22, self-aware voice at 1:47, comic argument at 2:27, and accent changes at 3:18.
- The most replayed 1:48.1–1:50.5 falls in the self-aware monologue, where expressive delivery is itself the proof. The second replay peak, 0:11.8–0:14.1, lands on the opening dialogue's first claim.
Plate 02 / Format card
| Platform · aspect · length | YouTube · 16:9 · 3:54 (source measured 640×360 at 30 fps) |
| Pace | One continuous motion-graphic canvas; 0 detected cuts; 172 internal picture changes, about 44/min. Conventional average shot length is 234 s and is misleading here. |
| Script | 475 transcribed words; 133 wpm over the full runtime, about 197 wpm while speaking; 62 short sentences averaging 7.7 words. |
| Voice | Ensemble of conversational, theatrical, comic, and storyteller archetypes; wide changes in energy and register. |
| On screen | Nearly 100% abstract gradient plus small typeset script; no visible presenter, footage, or product UI. |
| Captions | Script excerpt cards with speaker labels and small boxed delivery cues; not large word-by-word subtitles. |
| Sound | Soft tonal/cinematic bed beneath the voices (music-like, inferred); voice performances take priority; −19.3 LUFS integrated, −1.9 dBFS true peak. |
| Structure | Short brand reveal → eight performance chapters → quiet end card with product claim and URL. |
Pro / 13 plates + recipe
The first plates are open. The full dissection is Pro.
Get the timed shots, type and sound specs, and the runnable recipe.
Open the full blueprint- Plate 03 / The hook, frame by frame (first 9 seconds)Pro
- Plate 04 / Structure (beat sheet)Pro
- Plate 05 / Shot grammarPro
- Plate 06 / Visual style guidePro
- Plate 07 / Voice and deliveryPro
- Plate 08 / Sound designPro
- Plate 09 / PackagingPro
- Plate 10 / Script templatePro
- Plate 11 / Production planPro
- Plate 12 / PromptsPro
- Plate 13 / Edit recipePro
- Plate 14 / QA checklistPro
- Plate 15 / Make it yoursPro
- Recipe / runnable promptPro