Skip to content
refs
REF-2965Source onlineBlueprint rev 01

Voice-led canvas transformation on an iPad

A live tablet demo builds one scene through spoken edits, escalating from a rough sketch to a reflective sunset composition.

Maker @krea_aiFor KreaOriginal post ↗
Cut timeline · 4 shots00:37
00:0000:0500:1000:1500:2000:2500:3000:35

Field notes / why it works

The video starts at 0:00 with a spoken object request and a visibly rough pen sketch, then shows a recognizable result by about 0:04. Each later command changes the same composition, expanding from object details to a park at about 0:13 and a reflective final material around 0:32; the mostly continuous tablet view makes cause and effect easy to follow. Four shots across 37.4 seconds leave room to inspect each result without losing the real interaction.

01

Plate 01 / The format in one breath

A person speaks short creative instructions while drawing on a tablet. The same screen visibly changes from a rough sketch into a finished scene, then transforms again on fresh voice commands. The implicit promise is visible by about 0:04: words and marks can steer a result in real time.

Why it works

  • At 0:00, a pen draws a crude chair as a low-register voice names the intended object; the first generated chair is visible around 0:04. This gives a direct input-to-output proof.
  • The requests escalate from object to cushion (0:03.6), side table (0:07.5), park (0:10.8), sunset (0:13.2), flowers (0:18.5), and material change (about 0:28–0:32). Each new result raises the scope without resetting the scene.
  • The camera holds on the person's hand and tablet for most of the 37.4 s. That continuity makes the interaction legible; only three definite cuts interrupt it.
02

Plate 02 / Format card

Platform · aspect · lengthX · 16:9, 1280×720 · 37.4 s at 30 fps
Pace4 shots, 3 definite cuts = 4.8 cuts/min; 4.8–12.8 cuts/min allowing uncertain changes; 8.0 definite picture changes/min; mean shot 9.34 s
Script77 measured words, 147 measured wpm; 10 short statements averaging 7.7 words; no questions
VoiceOne conversational adult voice, low register, calm and matter-of-fact
On screenNearly all live-action, high oblique view of a person using a tablet; product output stays prominent
CaptionsSmall white sentence subtitles with dark outline, centered low; no kinetic highlight
SoundSpoken commands with a low tonal bed about 10 dB under voice; no confirmed accent effects
StructureInitial drawing → object result → scene additions → environment change → detail addition → material-change finale; no spoken CTA

Pro / 13 plates + recipe

The first plates are open. The full dissection is Pro.

Get the timed shots, type and sound specs, and the runnable recipe.

Open the full blueprint
  1. Plate 03 / The hook, frame by frame (first 3 seconds)Pro
  2. Plate 04 / Structure (beat sheet)Pro
  3. Plate 05 / Shot grammarPro
  4. Plate 06 / Visual style guidePro
  5. Plate 07 / Voice and deliveryPro
  6. Plate 08 / Sound designPro
  7. Plate 09 / PackagingPro
  8. Plate 10 / Script templatePro
  9. Plate 11 / Production planPro
  10. Plate 12 / PromptsPro
  11. Plate 13 / Edit recipePro
  12. Plate 14 / QA checklistPro
  13. Plate 15 / Make it yoursPro
  14. Recipe / runnable promptPro