Ship Platform Ready Verticals With an 8 Field Scene Breakdown Template

Use this AI-ready scene breakdown template to author scenes that plug directly into an automated episodic pipeline and preserve character continuity. The schema runs on eight required fields: scene_id, duration_seconds, hook_text (0 to 1 second), beat_summary, characters_with_continuity_IDs, anchor_frame_reference, ai_prompt_seed, and export_hints. Platforms like Iguanify accept this structure and turn it into produced, platform-ready vertical episodes without manual editing.
TL;DR:
- Using stored anchor images for characters at different angles and expressions is essential to prevent identity drift across scenes and episodes.
- The scene breakdown schema requires eight specific fields, including anchor references and timing tags, to ensure seamless automation and continuity.
- Scenes should be limited to 30 to 90 seconds, with the first second dedicated to a tightly written hook that captures viewers’ attention immediately.
- The AI prompt seed must follow a consistent pattern, specifying camera, lighting, anchor references, and aspect ratio to produce stable, platform-ready videos.
- Prioritize establishing reliable anchor references before refining prompts or visuals, as maintaining character identity is the foundation of effective automated episodic content.
Table of Contents
- Field-by-field template schema and formatting rules
- Quick-start example: a 45-second vertical scene
- Continuity and anchors: preventing drift across scenes
- AI prompt patterns and export rules for vertical scenes
- Hook-first writing, compression, and cliffhanger mapping
- Why this template fits an automated episodic workflow
- What actually matters in a scene breakdown template
- Turn the template into produced episodes
- FAQ
- Sources
- Selected research and primary sources
Field-by-field template schema and formatting rules
A scene breakdown built for automation is really a data contract between your writing and the engine that generates footage. Every field has to be readable by a human writer and parseable by a script. Here is the structure we recommend, with type, whether it’s mandatory, and the format an engine expects.
- scene_id: string, mandatory, format “EP01_SC03” so episode and scene order stay sortable.
- duration_seconds: integer, mandatory, 30 to 90 for most vertical micro-drama beats.
- hook_text: string, mandatory for scene 1 of any episode, the words or visual beat that lands in the first second.
- beat_summary: string, mandatory, one sentence describing the single emotional move in the scene.
- characters_with_continuity_IDs: array, mandatory, each character tagged with a stable ID like “char_mara_01” rather than a free-text name.
- anchor_frame_reference: string, mandatory once a character has appeared before, pointing to a stored anchor file.
- ai_prompt_seed: string, mandatory, the base prompt fragment the engine expands with camera and lighting tags.
- export_hints: object, mandatory, aspect ratio, caption safe zone, and target duration.
| Field | Type | Mandatory | Example |
|---|---|---|---|
| scene_id | string | yes | EP02_SC01 |
| duration_seconds | integer | yes | 45 |
| hook_text | string | scene 1 only | “She opens the letter she was told to burn.” |
| characters_with_continuity_IDs | array | yes | [“char_mara_01”] |
| anchor_frame_reference | string | if recurring | anchors/char_mara_01_front.png |
| ai_prompt_seed | string | yes | “medium shot, kitchen, tense pause” |
Name anchor files by character ID plus angle, so “char_mara_01_threequarter.png” stays unambiguous across a full season. Keep continuity IDs constant for a character even when their name changes mid-story, since the engine tracks identity by ID, not by dialogue. For the short-drama script format most vertical creators already use, this schema maps cleanly onto existing scene numbering.
Quick-start example: a 45-second vertical scene
Here’s a filled scene for a 45-second episode opener, annotated so you can see what each field triggers in an automated engine.
- scene_id: “EP01_SC01”, tells the pipeline this is the first scene of the first episode, which triggers full anchor generation rather than anchor reuse.
- duration_seconds: 45, sets the render length and determines how many shots the engine will cut between.
- hook_text: “The phone rings. She already knows who it is.”, lands in the first second and gets flagged for a tight opening shot with no establishing wide.
- characters_with_continuity_IDs: [“char_mara_01”], tells the engine to pull the stored anchor rather than generate a new face.
- anchor_frame_reference: “anchors/char_mara_01_front.png”, triggers anchor reuse instead of fresh character generation, which is what keeps her face identical to episode three.
- ai_prompt_seed: “close-up, dim kitchen light, hand hesitating over phone”, expands into camera, motion, and lighting instructions.
- export_hints: {aspect: “9:16”, duration: 45, caption_safe: true}, locks the output to vertical and reserves space for captions.
The cliffhanger gets encoded as a short tag inside beat_summary, something like “ends mid-ring, unresolved,” so the next scene’s engine pass knows to open on the same unanswered tension.
Continuity and anchors: preventing drift across scenes
Character drift is the most common failure mode in automated episodic production, and anchors are the fix researchers keep landing on. A multistage pipeline paper recommends moving from script to anchor creation to initial frames to shot generation to composition, treating the scene breakdown itself as the pipeline’s blueprint. Without anchors, identity consistency collapses across scenes; anchor-based pipelines keep it stable.
Separately, Gloria’s content-anchor research defines three anchor types worth storing for every recurring character:
- Global anchors: a clean front-facing reference image that defines the character’s base appearance.
- Viewpoint anchors: three-quarter and profile angles so the engine can rotate a character without losing identity.
- Expression anchors: a small bank of expressions, roughly eight, covering the emotional range the series needs.
Store anchors at a consistent resolution, name them by continuity ID plus angle or expression, and reference them directly inside anchor_frame_reference rather than describing them in prose.
Pro Tip: Build your expression bank once per character before writing episode one, not scene by scene, so every later episode draws from the same set.

AI prompt patterns and export rules for vertical scenes
Your ai_prompt_seed field should expand into a consistent shell every time, not a fresh invention per scene. A dependable pattern includes camera distance, motion verb, lighting descriptor, and an explicit anchor reference call, in that order, so the engine treats anchor usage as a parameter rather than a suggestion, as explained in Storyboard vs Shot List: What Every Production Needs.
- Camera and motion: specify shot type (close-up, medium, wide) and one motion cue (static, slow push, handheld).
- Lighting: one descriptor tied to mood (dim kitchen light, harsh noon sun, cold fluorescent).
- Anchor call: reference the stored file directly so the engine pulls identity rather than regenerating a face.
- Export aspect: 9:16 for nearly all serialized vertical formats, locked in export_hints.
- Target duration: 30 to 90 seconds per scene, consistent with YouTube’s guidance on Shorts, which allows vertical uploads up to three minutes.
- Caption margins: reserve the top and bottom 10 to 15% of frame for platform UI and captions.
Mark hooks and cliffhangers as their own metadata tags rather than burying them in narrative text, since that’s what lets an automated pass enforce precise timing on the opening second and the final unresolved beat.
Hook-first writing, compression, and cliffhanger mapping
Serialized vertical episodes live or die on the first second, so the template enforces that with a dedicated hook_text field capped tightly in length and placed only in scene one. A hook that takes two seconds to land is already too slow for how people scroll.
- Hook-first: write hook_text as a single visual or line that needs no setup, since setup is what causes drop-off.
- Compression: give each scene one emotional move, split roughly as hook, development, turn, and cliffhanger within your 30 to 90 second window.
- Cliffhanger mapping: tag the unresolved beat explicitly so the next episode’s opening scene can reference it directly.
- Series-level tension: track open threads in a simple metadata field per episode, not just inside dialogue.
Our cliffhanger idea bank and three-question test gives a fast way to check whether an ending beat is strong enough before it goes into production.
Why this template fits an automated episodic workflow
We built this schema around how an automated pipeline actually consumes a scene: structured fields, not loose prose. Iguanify accepts this format directly, uses stored anchors to keep characters consistent across episodes, and returns platform-ready vertical video with full ownership retained by the creator. For deeper workflow mapping, see our AI video production guide.

What actually matters in a scene breakdown template
Most advice on scene breakdowns for AI video treats the template as a formatting exercise, filling in boxes so a script “looks organized.” That misses the real function. A scene breakdown for an automated pipeline is a continuity contract, and the fields that matter most are the ones an engine can act on without a human reinterpreting them: anchor references, continuity IDs, and explicit timing tags.
Where conventional advice falls short is treating hooks and cliffhangers as creative flourishes rather than data. If your hook lives only in the dialogue and not in a dedicated field, nothing forces the engine to prioritize it in the first second of render. The same goes for anchors. A character description in prose is not an anchor. A stored, named image file referenced by ID is.
Prioritize the anchor system before you worry about prompt phrasing or export polish. Get continuity right first. Everything else, camera language, lighting tags, caption margins, is easier to fix after the fact than a face that drifts between episode two and episode three.
— Leonard
Turn the template into produced episodes
Once your scene breakdown is filled in, the template itself is just the blueprint. We built Iguanify to take that blueprint and produce the episode: your anchors keep characters consistent from scene to scene and episode to episode, your prompt seeds expand into full shots, and the output arrives sized and captioned for TikTok, YouTube, or Instagram, with full ownership staying with you.

If you want to try it yourself, export your scene breakdown as JSON and run it through the AI drama generator; current prices are on the pricing page. Prefer a hands-off route where we produce the series for you? Start from our main production page and describe your premise, and we’ll take it from there.
FAQ
What fields does an AI-ready scene breakdown template need?
At minimum, you need scene_id, duration_seconds, hook_text, beat_summary, characters_with_continuity_IDs, anchor_frame_reference, ai_prompt_seed, and export_hints. These eight fields let an automated pipeline parse your scene without a human filling gaps.
How long should each scene be in a vertical micro-drama?
Most vertical episode scenes run 30 to 90 seconds, built around one emotional turn per scene. Full episodes can run up to three minutes under YouTube’s Shorts classification for vertical uploads.
What is an anchor frame and why does it matter?
An anchor frame is a stored reference image, global, viewpoint, or expression, that an automated engine pulls instead of regenerating a character’s face from scratch. Content-anchor research shows this is what keeps identity stable across scenes and episodes.
Can Iguanify use a scene breakdown template I built myself?
Yes, Iguanify accepts structured scene breakdowns and uses the anchor and continuity fields to keep characters consistent across a full series. Pricing for a Standard episode or Pilot is $34.99 one-off per episode through the AI drama generator.
How do I prevent character drift across episodes?
Store a global, viewpoint, and expression anchor for each recurring character and reference the same file by continuity ID in every scene that uses them. A multistage pipeline approach, moving from script to anchors to initial frames to shot generation, keeps identity stable where ad hoc generation tends to drift.
Sources
- Multi-Shot Character Consistency for Text-to-Video Generation (arXiv 2412.07750v1)
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories (arXiv 2512.16954)
- Gloria: Consistent Character Video Generation via Content Anchors (CVPR 2026)
- Understand three-minute YouTube Shorts - YouTube Help
Selected research and primary sources
- Multi-Shot Character Consistency for Text-to-Video Generation
- Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
- Gloria: Consistent Character Video Generation via Content Anchors
- Understand three-minute YouTube Shorts
IGUANIFY