Transformer, From Word Puzzle to Attention Math
A six-minute animated Chinese lesson that moves from a one-word puzzle to a worked attention calculation on one continuous dark canvas.
Field notes / why it works
It names a hidden mechanism behind familiar apps at 0:08, then uses a one-word sentence change at 0:31–0:58 to make context feel necessary. A three-stage visual map at 1:51 prepares viewers for the concrete score, scale, and softmax calculation at 4:50–5:59; 36 major graphic changes sustain a single uncut canvas.
Plate 01 / The format in one breath
A fast but patient narrated lesson starts with familiar applications, promises an explanation accessible to students, then turns an everyday language puzzle into a reason to learn increasingly precise machinery. A single dark canvas changes diagrams in place while the narration moves from intuition to vectors, then to a numerical attention calculation.
Why it works
- Familiar AI applications lead to the named concept at 0:08; the viewer gets a reason to care before any notation.
- At 0:31–0:58, changing one word flips the answer to a pronoun puzzle. That concrete contradiction creates the need for context and an open loop.
- At 1:51, a three-step map gives orientation before the lesson gets mathematical; at 4:50–5:59, the viewer can follow actual scores, scaling, and probabilities.
Plate 02 / Format card
| Platform · aspect · length | X · 1280×720, 16:9, 30 fps · available asset 360.1 s |
| Pace | One continuous designed canvas, 0 hard cuts, average shot 360.1 s; 36 detected major internal steps, 6/min. Small icon builds begin at 0:00.75 and were missed by the detector. |
| Script | Chinese narration; 1,246 machine-segmented “words,” about 208 measured units/min overall and 248 while speaking. These units are not directly comparable with English word counts. Three detected questions; transcript sentence segmentation is unreliable. |
| Voice | Bright, high-register, animated explainer; authoritative but approachable. No person appears. |
| On screen | Nearly 100% code-like motion graphics: icons, labelled cards, token chips, matrices, equations, network diagrams. No live action or stock footage observed. |
| Captions | Small white, centre-aligned Chinese subtitles in a translucent near-black strip at the bottom; key phrases also appear as larger coloured labels in the diagram. |
| Sound | Continuous soft tonal bed, roughly 19 dB under speech; -16.0 LUFS integrated, -1.5 dBFS true peak. Discrete effects are not established by the supplied evidence. |
| Structure | Applications → promise → one-word puzzle → older method → three-step map → token/vector intuition → dot-product primer → Q/K/V analogy → worked attention calculation. The measured excerpt ends during the calculation, with no observed CTA. |
Pro / 13 plates + recipe
The first plates are open. The full dissection is Pro.
Get the timed shots, type and sound specs, and the runnable recipe.
Open the full blueprint- Plate 03 / The hook, frame by frame (first 10 seconds)Pro
- Plate 04 / Structure (beat sheet)Pro
- Plate 05 / Shot grammarPro
- Plate 06 / Visual style guidePro
- Plate 07 / Voice and deliveryPro
- Plate 08 / Sound designPro
- Plate 09 / PackagingPro
- Plate 10 / Script templatePro
- Plate 11 / Production planPro
- Plate 12 / PromptsPro
- Plate 13 / Edit recipePro
- Plate 14 / QA checklistPro
- Plate 15 / Make it yoursPro
- Recipe / runnable promptPro