Skip to content
refs
REF-0106Source onlineBlueprint rev 01

Transformer, From Word Puzzle to Attention Math

A six-minute animated Chinese lesson that moves from a one-word puzzle to a worked attention calculation on one continuous dark canvas.

Maker @doteyOriginal post ↗
Cut timeline · 1 shots in the first 06:0012:12
00:0002:0004:0006:0008:0010:0012:00
Measured first 06:00

Field notes / why it works

It names a hidden mechanism behind familiar apps at 0:08, then uses a one-word sentence change at 0:31–0:58 to make context feel necessary. A three-stage visual map at 1:51 prepares viewers for the concrete score, scale, and softmax calculation at 4:50–5:59; 36 major graphic changes sustain a single uncut canvas.

01

Plate 01 / The format in one breath

A fast but patient narrated lesson starts with familiar applications, promises an explanation accessible to students, then turns an everyday language puzzle into a reason to learn increasingly precise machinery. A single dark canvas changes diagrams in place while the narration moves from intuition to vectors, then to a numerical attention calculation.

Why it works

  • Familiar AI applications lead to the named concept at 0:08; the viewer gets a reason to care before any notation.
  • At 0:31–0:58, changing one word flips the answer to a pronoun puzzle. That concrete contradiction creates the need for context and an open loop.
  • At 1:51, a three-step map gives orientation before the lesson gets mathematical; at 4:50–5:59, the viewer can follow actual scores, scaling, and probabilities.
02

Plate 02 / Format card

Platform · aspect · lengthX · 1280×720, 16:9, 30 fps · available asset 360.1 s
PaceOne continuous designed canvas, 0 hard cuts, average shot 360.1 s; 36 detected major internal steps, 6/min. Small icon builds begin at 0:00.75 and were missed by the detector.
ScriptChinese narration; 1,246 machine-segmented “words,” about 208 measured units/min overall and 248 while speaking. These units are not directly comparable with English word counts. Three detected questions; transcript sentence segmentation is unreliable.
VoiceBright, high-register, animated explainer; authoritative but approachable. No person appears.
On screenNearly 100% code-like motion graphics: icons, labelled cards, token chips, matrices, equations, network diagrams. No live action or stock footage observed.
CaptionsSmall white, centre-aligned Chinese subtitles in a translucent near-black strip at the bottom; key phrases also appear as larger coloured labels in the diagram.
SoundContinuous soft tonal bed, roughly 19 dB under speech; -16.0 LUFS integrated, -1.5 dBFS true peak. Discrete effects are not established by the supplied evidence.
StructureApplications → promise → one-word puzzle → older method → three-step map → token/vector intuition → dot-product primer → Q/K/V analogy → worked attention calculation. The measured excerpt ends during the calculation, with no observed CTA.

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 10 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro