Skip to content
refs
REF-0928Source onlineBlueprint rev 01

Veo 3.1's Three Start-to-End Transformations

A 40-second feature film uses paired endpoint thumbnails and three cinematic AI-generated bridges to show frame-guided video control.

Maker Google DeepMindFor Google DeepMind / VeoOriginal post ↗
Cut timeline · 4 shots00:40
00:0000:0500:1000:1500:2000:2500:3000:3500:40

Field notes / why it works

The first five seconds pair a concise control promise with a labeled barn start frame, so the viewer understands the mechanism before the first cut at 12.9 seconds. Two more visually distinct transformations at 12.9 and 22.7 seconds refresh attention while endpoint thumbnails make each result verifiable; the final 3.9 seconds give the brand a clean resolution.

01

Plate 01 / The format in one breath

A brief, low-register narrator explains that the viewer can specify two endpoints; three cinematic examples then demonstrate the transformation with almost no further speech. Paired start/end thumbnail overlays make the mechanism legible while the generated motion supplies the spectacle. The final four seconds resolve to a restrained company end card.

Why it works

  • The opening states the control benefit before the visual payoff: by 0:04.5, the narrator has named start and end points while a barn shot carries a centered “First and last frame” label.
  • A long 0:00–0:12.9 first example has room to develop: barn interior/exterior imagery yields to a sunlit horse rider in a field, with the paired endpoint thumbnails visible around 0:03.2 and 0:06.4.
  • Two more subjects reset attention without more explanation: an abstract faceted figure becomes an athletic pose around 0:16–0:22.7, then a cowgirl close-up leads toward a cactus/flower image around 0:26–0:35.9.
02

Plate 02 / Format card

Platform · aspect · lengthYouTube · 16:9 · 640×360 source · 39.8 s · 30 fps
Pace4 measured shots, 3 cuts, 4.5 cuts/min; 4 changes inside shots, 10.5 total picture changes/min; average shot 9.96 s
Script19 transcribed words; 148 wpm across the speech span, 175 wpm while speaking; two short sentences; speech covers 19% of run time
VoiceOne low-register, measured, confident adult narrator; no on-camera speaker
On screenAbout 90% cinematic transformation footage and composited endpoint thumbnails; about 10% dark brand card
CaptionsNo speech subtitles; one small centered white feature label in the opening, plus image-pair overlays
SoundContinuous music-like bed, comparatively forward in the mix; brief narration only at the start; no clearly evidenced cut effects
StructurePromise and example 1 → example 2 → example 3 → brand stop; visual payoffs throughout, no spoken CTA

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 3 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro