Skip to content
refs
REF-0635Source onlineBlueprint rev 01

A robot proves speech reasoning at the table

A lab-style capability announcement lets a long, visible robot task and natural dialogue provide the proof.

Maker FigureFor FigureOriginal post ↗
Cut timeline · 11 shots02:35
00:0000:3001:0001:3002:0002:30

Field notes / why it works

The capability is labeled within the first 1.3 seconds, but the film waits until 0:15.4 for the first spoken observation, creating anticipation around the robot close-up; the top replay peak is 0:13.9–0:15.5. A 73.66-second wide master shot from 0:28 shows the tasks and their outcomes in one workspace, giving the spoken reasoning visible stakes. Close hand details at 2:07.8–2:28.2 summarize the dexterity after the uninterrupted proof.

01

Plate 01 / The format in one breath

A clinical status card labels a capability, then a real person asks a machine to perceive a tabletop, act on objects, explain its choices, and evaluate the result. The promise is made in text by 0:01.3; the demonstration earns it in a mostly locked-off take rather than rapid editing.

Why it works

  • At 0:00–0:03, date, capability label, and geometric mark build a technical frame before any voice; the first picture change occurs at 0:00.75.
  • The 0:28–1:41.7 master shot stays on the workspace for 73.66 seconds, so the viewer can follow the cause and effect of each instruction and movement.
  • The most-replayed point, 0:13.9–0:15.5, coincides with the robot close-up and first spoken observation; later replay peaks at 0:34–0:35.6 and 0:49.6–0:51.1 fall near the first request/action and its explanation.
02

Plate 02 / Format card

Platform · aspect · lengthYouTube · 16:9 · 1280×720 · 23.98 fps · 2:34.7
Pace11 shots; 10 cuts (3.9/min); 38 measured within-shot changes; 18.6 picture changes/min including cuts. Mean shot 14.06 s; median 6.48 s; longest 73.66 s. Some measured steps appear to be object movement, not overlays (inferred from sheets).
Script179 measured words; 98 wpm across full runtime, 218 wpm during speech. 13 sentences; four questions. First speech at 0:15.4.
VoiceTwo conversational adult archetypes: a low, smooth synthetic-sounding assistant and a casually directive human tester (inferred from transcript and frames).
On screenAbout 14% opening status/close-up; about 73% wide or medium live action proof; about 8% action detail inserts; about 4% black end card (estimated from shot boundaries).
CaptionsNo running subtitles. A single uppercase capability claim appears over the final detail montage at 2:07.8–2:28.2.
SoundNear-silent room tone before speech and between exchanges; no bed under dialogue; music-like outro from about 2:04.6 (measurement).
StructureStatus/promise → scene read → request and object handling → explanation → second task → self-assessment → brief proof montage → brand end card. No spoken CTA.

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 3 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro