Skip to content
refs
REF-0523Source onlineBlueprint rev 01

Glowing Orbs Become an AI Creation Flow

A 24-second SaaS motion showcase turns luminous capability orbs into a sparse creation interface and a finished runner ad.

Maker @ankush_motionFor ElevenLabsOriginal post ↗
Cut timeline · 3 shots00:24
00:0000:0500:1000:1500:20

Field notes / why it works

Seven bright orbs collapse and reform into labelled capability chips over the first 2.75 seconds, creating movement without an opening line. A hard black-to-white change at 5.2 seconds reveals the creation interface, then the runner output and setting chips at 18.2–20.9 seconds provide a concrete payoff. The three long shots let each stage develop while animation supplies motion inside them.

01

Plate 01 / The format in one breath

A polished SaaS motion sample presents a product through transformation: seven luminous objects gather on black, turn into small labelled capabilities, then yield to a spare white creation interface and a finished media asset. Its promise becomes legible through the workflow around 0:08–0:14, with the generated result held from about 0:18 onward.

Why it works

  • The first 2.75 s compress seven saturated orbs into a single mark, creating motion and a visual question without copy or narration.
  • The tonal flip from black to white at 0:05.2 resets attention; the sparse prompt field at 0:08.6 makes the abstract opening feel like a product workflow.
  • The last third shows an actual runner image and language/voice controls (0:18.2–0:20.9), giving the product sequence a concrete output.
02

Plate 02 / Format card

Platform · aspect · lengthX · 16:9 · 1280×720 · 30 fps · 24.0 s
Pace3 measured shots, 2 definite cuts, 5–10 cuts/min counting two uncertain changes; average shot 8.0 s; 1 detected internal step change
Script22 machine-transcribed words in a short late passage; text is mainly interface copy. Transcript confidence is low, so use semantic beats, not its wording.
VoiceA brief, fast spoken passage is detected from 0:16.4; apparent 455 wpm is unreliable because 14/22 words have low confidence.
On screenAbout 22% dark orb animation · 43% white interface/process · 35% output showcase
CaptionsNo conventional subtitles visible; use small native interface labels and short status text.
SoundQuiet continuous bed inferred from audio analysis; rough tempo 110 BPM, low confidence. Integrated −22.6 LUFS, true peak −0.4 dBFS.
Structure3 visual acts: abstraction → creation flow → finished output; payoff from ~0:18, no visible in-video CTA.
03

Plate 03 / The hook, frame by frame (first 3 seconds)

TimePictureOn-screen textWords and soundWhat it does
0:00Seven glowing pastel spheres across central black fieldNoneNo measured speechImmediate visual pattern interrupt.
0:00.25Spheres converge diagonally toward centreNoneBed continues (inferred)Starts a question: what are these parts becoming?
0:00.50–0:01.00One tiny merged iridescent form hovers at centreNoneNo measured speechContrast in scale creates a pause.
0:01.25–0:01.75A vertical set of orb-backed glass labels emerges; legible examples include “Deep,” “Clarity,” and “Lame”Small white sans labelsNo measured speechTurns colour into capability tokens.
0:02.00–0:02.75Labels rearrange; “Calm,” “Deep,” and “Clarity” stack, then collapse into beadsWhite text on thin outlined cardsNo measured speechGives the abstraction a product-like taxonomy.

Hook formula: visual pattern interrupt + open loop: {{N}} vivid {{CAPABILITY_SYMBOLS}} merge into one {{PRODUCT_MARK}} before the product interface appears.

Why it holds: There are no cuts in the first 3 s, yet the objects change shape and spacing at quarter-second intervals. The viewer waits to see what the glowing pieces represent.

04

Plate 04 / Structure (beat sheet)

#TimeBeatWhat happensTechniqueTemplate line
10:00–0:05.2Abstract hookColour orbs merge, become short labelled traits, and reorganize.Black negative space, soft bloom, particle motion.{{CAPABILITIES}} converge into {{PRODUCT_MARK}}.
20:05.2–0:15.4Product processCut to a white “Flows Agent” style prompt, then enlarged text and a status/process panel; by 0:13.5 the runner concept appears beside the agent state.High-contrast palette switch; editorial crops of a clean interface.Enter {{TASK_PROMPT}}; show {{PROCESS_STATUS}}.
30:15.4–0:24.0Output payoffGenerated runner scene fills a card; language selector and voice chip become visible around 0:19.7–0:20.9.Hold the result long enough to inspect; add compact controls beneath.Reveal {{FINISHED_OUTPUT}} with {{SETTING_1}} and {{SETTING_2}}.

Retention map: visible shifts at 0:05.2, ~0:13.2 (uncertain cut), 0:15.4, ~0:18.0 (uncertain cut), and 0:20.6 (internal step) · payoff ~0:18.2 · in-video CTA none visible · ending appears to hold/stop rather than loop (inferred from supplied frames).

05

Plate 05 / Shot grammar

Mix: Designed orb animation ~22%; interface motion / screen-like composition ~43%; generated-media showcase ~35%. No presenter or talking head appears.

Rhythm: The measured edit has two definite hard scene changes (0:05.2, 0:15.4). Animation within each long shot supplies most of the perceived pace. At 0:13.2 and 0:18.0 the detector is uncertain, so treat them as compositional shifts rather than confirmed cuts.

Rules inferred from the sheets:

  • Begin on a nearly empty black canvas so each bright object is conspicuous.
  • Switch to white exactly when the story moves from metaphor to tool.
  • Keep interface surfaces shallow and uncluttered: one prompt, one process trail, one output card.
  • Allow the proof shot to breathe for several seconds instead of cutting away immediately.

Full shot list

#InDurTypeFramingWhat's on screenText / overlayInto the next shot
10:00.05.17 s2D/3D motion graphicWide, centred on blackGlowing spheres and stacked trait chipsShort labels such as “Deep,” “Clarity,” “Calm”Hard cut / palette flip to white
20:05.210.23 sInterface motion graphicWide UI, then extreme crop, then compositePrompt field, oversized text fragment, process/status panel, runner previewSmall dark UI copy, green indicator, orange process emphasisDefinite cut near 0:15.4; possible internal shift at 0:13.2
30:15.48.6 sOutput/demo compositionLarge centred landscape media cardRunner footage/image, then language and voice controls“English” chip; named voice chip in sourceHolds toward 0:24; step at 0:20.6
06

Plate 06 / Visual style guide

  • Look: Two-stage palette: near black #030303 and glowing lavender #C9A9FF, cyan #35C7F4, coral #FF6A6B, warm yellow #FFE680; then white #FFFFFF, charcoal text #1B1B1B, pale grey fields #F1F1F1, sparse green state dot #3DE087. The final runner asset adds teal sky and warm orange ground.
  • Interface type: Closest free match is Inter, regular/medium for tiny UI copy and semibold for labels. Prompt and status text look roughly 12–20 px in a 1280×720 frame; the momentary cropped typography at 0:10.3 is deliberately huge.
  • Captions: None in the usual subtitle sense. Put only 1–3 word capability labels in slim rounded rectangles; use white text over black in the opening, dark text over white in the UI. Avoid karaoke captions.
  • Graphics and overlays: Spheres have soft coloured internal gradients, bloom, and specular shading. Chips have hairline white/grey borders, 8–12 px rounded corners, and subtle shadows. The UI keeps generous empty margins.
  • Media card: Rounded landscape frame, about one third to one half of the screen width in the final section; put compact pill controls below it.
  • Transitions: Orb motion is continuous; the main act changes are hard cuts or rapid graphic reveals. No visible whip, glitch, or camera shake effect is required.
  • Branding: Small agent label and status indicator convey the product. No prominent end-card logo is visible in supplied sheets.
07

Plate 07 / Voice and delivery

Measured: Automatic transcription detects 22 words beginning 0:16.36; speech covers roughly 12% of the video. It reports 455 wpm, no ≥0.25 s pauses, low median pitch, and wide pitch movement. Fourteen words are low confidence, so those detailed voice metrics and exact words are uncertain.

Described: Use a concise, bright, capable product-demo narrator for the output reveal, or use one short synthetic voice sample. Keep it subordinate to the visual proof. The source passage is too uncertain to justify an accent, age, or exact cadence claim.

Voice direction:

Deliver one or two short product-result lines in a clear, energetic conversational voice. Start only when the finished output appears. Keep diction crisp and the line under four seconds; leave the abstract hook and input setup free of narration.

08

Plate 08 / Sound design

  • Music: A low-level electronic/ambient pulse seems present throughout (inferred). Analysis suggests ~110 BPM, but confidence is only 0.22; choose by feel rather than matching that number.
  • Sound effects: Gentle orb convergence swells, soft UI clicks for chips, and one restrained lift into the final output are appropriate production choices (inferred, not verified source effects).
  • Silence and space: No measured speech until 0:16.4. Preserve this long visual-first opening.
  • Mix targets: Source measures −22.6 LUFS integrated and −0.4 dBFS true peak. For a new export, target roughly −16 to −14 LUFS and ≤−1 dBTP after adding a reliable narration and music mix; this is a production recommendation, not a source measurement.
09

Plate 09 / Packaging

  • Title: Format is a short category/showcase claim: {{PRODUCT}}: from {{INPUT}} to {{OUTPUT}}.
  • Cover / thumbnail: Seven luminous multicolour beads in a loose vertical arrangement on pitch black, centred with large margins; no thumbnail text or face.
  • Description, hashtags, pinned comment: Source post says it is a SaaS motion project and includes a commission invitation. No hashtag or pinned-comment pattern is documented.
  • Audience signals: 5,943 views, 200 likes, 19 comments, 6 reposts at capture, roughly 33.7 likes and 3.2 comments per 1,000 views. No most-replayed data or comment text is supplied, so no audience reaction can be assigned to a specific moment.
10

Plate 10 / Script template

This is principally a visual script. Keep spoken copy to about 8–18 words over the final output; on-screen interface text carries the middle act.

[0:00–0:05.2 · ABSTRACT HOOK · 0 spoken words]
Show {{N}} glowing {{CAPABILITY_SYMBOLS}} combine into {{PRODUCT_MARK}}. Label only 2–4 traits: {{TRAIT_1}}, {{TRAIT_2}}, {{TRAIT_3}}.

[0:05.2–0:15.4 · PROCESS · 0 spoken words]
In a minimal {{PRODUCT}} interface, enter a one-line request: “{{TASK_PROMPT}}”. Show 2–3 short status updates: {{STATUS_1}}, {{STATUS_2}}, {{STATUS_3}}. Introduce a preview of {{OUTPUT}}.

[0:15.4–0:24.0 · RESULT · ~8–18 spoken words]
Hold the finished {{OUTPUT}}. Reveal {{SETTING_1}} and {{SETTING_2}} as small chips. Voice: “{{PRODUCT}} turns {{INPUT}} into {{BENEFIT}}.”
11

Plate 11 / Production plan

BeatTimeVisualOn-screen textAudioNotes
Orb hook0:00–0:05.27 spheres, converge and resolve into trait chips2–4 short trait labelsAmbient pulse, optional soft swellAnimate position and scale continuously; keep black clear.
Interface0:05.2–0:15.4White prompt field → cropped type → process panel and previewPrompt + 2–3 status linesMinimal UI ticksAvoid dense readable paragraphs; give each state ~1–3 s.
Result0:15.4–0:24.0Large finished media card; settings appear beneath1 language/format chip and 1 voice/style chipBrief narration + gentle riseUse an original or licensed hero asset.
12

Plate 12 / Prompts

Script prompt

Write a new 24-second SaaS motion-film script for {{PRODUCT}}, serving {{AUDIENCE}}, that transforms {{INPUT}} into {{OUTPUT}}. Use three timed acts: 0–5.2 s abstract capability orbs with no speech; 5.2–15.4 s clean white product interface showing one user request and three terse process states; 15.4–24 s finished output plus two small setting chips. Write original interface copy, 2–4 one-word capability labels, and at most 18 spoken words during the last act. Make the benefit legible without claiming unsupported features. Return frame-by-frame visual and audio columns. Avoid referencing any real creator, brand voice, or source script.

Voice prompt

Clear, bright, concise product-demo narrator. Conversational energy, crisp consonants, one compact sentence only during the result shot, no celebrity or identifiable-person imitation.

Visual prompts

  • Orb field: Seven small iridescent spheres suspended on pure black, lavender cyan coral and buttery yellow internal gradients, soft bloom and clean silhouettes, centred with generous negative space, no text, no logo, 16:9.
  • Interface: Minimal white SaaS creation interface for {{PRODUCT}}, one light grey rounded prompt input containing {{TASK_PROMPT}}, tiny dark Inter labels, subtle green status dot, abundant white space, 16:9.
  • Result: Original {{OUTPUT}} for {{TOPIC}}, vivid professional hero image or footage in a softly rounded landscape card on white, with two understated setting pills below, 16:9.

Caption spec: No subtitles unless accessibility requires them. UI labels: Inter Medium 18–24 px for chips, dark #1B1B1B on white or white on black, short fades or 100–150 ms scale pops. Keep spoken-word captions separate from product UI if added.

Music prompt: Sparse warm electronic ambient beat, ~100–115 BPM, gentle plucks and low soft pulse, no aggressive percussion; small lift as the result appears.

13

Plate 13 / Edit recipe

  1. Lock a 24 s 1280×720, 30 fps timeline and choose the final proof asset first.
  2. Build 0–5.2 s on black: seven coloured orbs converge, resolve into a compact stack of capability chips, then disperse.
  3. Hard cut to white at 5.2 s; animate a single prompt entry, one oversized text crop, and process/status copy through 15.4 s.
  4. Reveal the actual {{OUTPUT}} at 15.4 s, refine/replace the hero image around 18 s, and add setting chips around 19.7–20.9 s.
  5. Record or synthesize a short original line only for the result, add quiet pulse music, then mix and export.
14

Plate 14 / QA checklist

  • Seven distinct objects are visible immediately, with meaningful convergence by 0:00.5.
  • The metaphor resolves into capability labels by ~0:01.5.
  • White product UI appears by 0:05.2, with a readable task by ~0:08.6.
  • A clear output preview is visible by ~0:13.5 and the finished asset is prominent by ~0:18.2.
  • Only 2–4 short trait labels and 2–3 process states compete for attention.
  • Narration, if used, starts with the reveal and is intelligible; do not use the source transcript as copy.
  • Export is 1280×720 at 30 fps; audio peaks below −1 dBTP.
15

Plate 15 / Make it yours

  • The creator's signature: The particular orb shader, named interface, runner footage, voice label, and client identity belong to this execution.
  • The transferable format: Abstract capability symbols → a sparse request/process UI → an inspectable finished asset, with a hard black-to-white pivot and minimal late narration.
  • Use your own product names, interface, output media, voice, and licensed sound.

<sub>Source: https://x.com/ankush_motion/status/2106444826539606152 · dissect 1.1.0 · 2026-10-07 · statements labelled “inferred” are judgment calls</sub>

Recipe / REF-0523 / rev 01
1,278 chars
Make an original 24-second, 1280×720, 30 fps SaaS motion film for {{PRODUCT}} about {{TOPIC}}, aimed at {{AUDIENCE}}. Use Remotion for animation and rendering, then ffmpeg for the final audio mix and H.264 export. Structure it in three acts. From 0–5.2 seconds, place seven small glowing, iridescent objects on near-black; converge them into one mark, expand into three thin outlined cards labelled {{CAPABILITY_1}}, {{CAPABILITY_2}}, and {{CAPABILITY_3}}, then collapse them again. From 5.2–15.4 seconds, hard cut to a white, spacious {{PRODUCT}} interface. Enter the original request {{TASK_PROMPT}}, pass through two or three short status lines, use one dramatic oversized text crop, and show a first preview of {{OUTPUT}}. From 15.4–24 seconds, show the finished {{OUTPUT}} in a large rounded media card; reveal {{SETTING_1}} and {{SETTING_2}} as compact pills between 19.7 and 20.9 seconds. Use Inter, hairline borders, black text, soft grey fields, and lavender/cyan/coral/yellow gradients only in the opening. Keep motion smooth, with the black-to-white cut as the main transition. Add a subdued warm electronic pulse and one short original spoken line at the result reveal; keep it clear above music. Use only original or licensed media, and finish without a forced CTA.