Skip to content
refs
REF-1238Source onlineBlueprint rev 01

A live conversation reveals the new model

A fixed two-view phone conversation demonstrates visual awareness before revealing the product announcement.

Maker @OpenAIFor OpenAIOriginal post ↗
Cut timeline · 2 shots01:22
00:0000:1500:3000:4501:0001:15

Field notes / why it works

The two-view frame makes the exchange legible from 0:00, and specific observations about the person and room at 0:05–0:13 turn the product's claim into visible evidence. A question at 0:14 leads to a detailed answer through 0:34; the identity hint at 0:50 delays the explicit capability reveal until 1:12. With only one measured cut across 82 seconds, the interaction itself carries attention.

01

Plate 01 / The format in one breath

A person casually talks with a product through a phone. Two simultaneous views let the audience see both the person and the device, while a sequence of questions makes the product's perception and conversational ability the evidence. The explicit announcement arrives only after the product has already responded to the room and the person.

Why it works

  • The 0:00 frame shows the whole experiment: participant on the left, live phone view on the right. There is no title card to delay the proof.
  • The exchange progresses from greeting (0:00–0:05) to observations about the participant and room (0:05–0:13), then a guess about the setup (0:20–0:34). Those answers make the later claim credible.
  • At 0:50–0:55 the participant suggests that the conversation partner itself is the announcement; the product statement follows at 1:12. The reveal comes from the interaction, not a graphic.
02

Plate 02 / Format card

Platform · aspect · lengthYouTube · 16:9 · 82.0 s · source measured at 640×360, 29.97 fps
Pace2 measured shots, 1 cut at 0:17.48 · 0.7 cuts/min · mean shot 41.02 s; framing stays almost identical across the cut
Script182 transcribed words, 135 wpm over runtime, 229 wpm during detected speech · 16 sentences averaging 11.4 words · 7 questions · about 4.9 “you” per 100 words. Transcript has a 0:55–1:12 gap; treat its word count as incomplete.
VoiceRelaxed adult participant and quick, upbeat conversational response voice; both medium register (inferred from dialogue and measured composite pitch)
On screen100% live-action, fixed two-view composite: about 69% medium shot of participant, 31% device close-up
CaptionsNone. No headlines, lower thirds, callouts or animated type in the frames.
SoundDialogue and device response carry the scene. Analysis detects a faint tonal bed, but cannot establish whether it is music or device/room audio. No clear cut effects.
Structure6 dialogue beats · identity hint at 0:50 · product capability reveal at 1:12–1:22 · no in-video CTA

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 3 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro