Skip to content
refs
REF-2079Source onlineBlueprint rev 01

Make Your Avatar, Then Meet Your Cartoon Self

A music-led avatar feature launch that turns editor controls into rapid real-person/cartoon reveals.

Maker @duolingoFor duolingoOriginal post ↗
Cut timeline · 24 shots00:30
00:0000:0500:1000:1500:2000:2500:30

Field notes / why it works

A phone notification makes the invitation legible within the first half-second, then 0:07.7–0:16.6 shows real customization controls beside people using phones. The edit accelerates into roughly half-second person/avatar matches at 0:18.4–0:23.0; the 0:22.3–0:22.6 match is the most-replayed interval. A varied character grid leads to a clear create-and-share card held for the last three seconds.

01

Plate 01 / The format in one breath

A word-light feature announcement turns a personalized app action into a social reveal. A notification creates the question, quick glimpses show where to make something, then real people are paired with illustrated versions of themselves in a rapid payoff montage. The final card tells viewers to make and share their own.

Why it works

  • The opening notification is readable by 0:00.5, so the viewer knows the action before seeing the interface (0:00–0:03).
  • The 0:07.7–0:16.6 middle shows an actual editing path—people holding phones beside avatar controls—so the promised result feels achievable.
  • Ten half-second transformation/reaction shots from 0:18.4–0:23.0 create the fastest section; the most-replayed interval, 0:22.3–0:22.6, is a person/avatar pairing.
02

Plate 02 / Format card

Platform · aspect · lengthYouTube · 16:9 · 30.0 s · measured source 640×360 at 23.98 fps
Pace24 shots, 23 cuts; 46–50 cuts/min allowing two uncertain changes; 50 picture changes/min; mean 1.25 s, median 0.95 s
ScriptAlmost entirely on-screen copy: one notification, one instruction card, one final action card. The five detected transcript words are low-confidence music transcription, so no reliable spoken-word rate.
VoiceNo confirmed narrator or intelligible dialogue; visual demonstration carries the message.
On screenAbout 45% live-action people/phone plates, 40% avatar or interface composites, 15% instruction/end cards (estimated from shot timings).
CaptionsNo speech captions. Short centered green display text on white at 0:05.8 and 0:27.1.
SoundContinuous tonal, upbeat bed at roughly 125 BPM (instrumentation inferred); −11.5 LUFS integrated, −0.3 dBFS true peak; no reliable cut-hit pattern.
StructureNotification hook → entry instruction → interface proof → avatar/reaction payoff → many-result grid → create/share CTA.

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 3 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro