---
name: paper-diorama-producer
description: >-
  Turn a script, fact list, or idea into a finished multi-minute layered
  cut-paper diorama film — flat matte cardstock worlds in a strict three-colour
  palette, a tiny figure in a huge symmetrical set, action-dense scenes joined
  by paper wipes, native impact SFX, ending on a face. End-to-end and
  SELF-CONTAINED: palette maths, exposure spec, wipe grammar, prompt templates,
  QC scripts, and the whole pipeline live in one folder — it never depends on
  another skill. Runs on Unsora; any voiceover is supplied or from a session TTS
  tool, then mixed with ffmpeg. Use whenever the user wants a paper-craft /
  cut-paper / papercut / diorama video longer than ~15 seconds, an end-to-end
  diorama film from a script, a diorama explainer or narrative short, wants to
  turn an idea into a multi-scene diorama animation, wants one diorama shot's
  hero-frame + Seedance prompts, or asks why a paper-craft clip feels empty,
  flat, or off its palette. Default to this for any multi-minute paper-diorama
  film on Unsora.
---

# Paper Diorama Producer

Script in, finished layered cut-paper diorama film out — minutes long, many
shots, wipe-joined — or a single shot's prompts, if that's all that's asked.
Self-contained: the visual system, the palette maths, the prompt templates, the
QC scripts, and the production pipeline all live in this folder. It **never**
refers to another skill.

**Hard precondition:** Unsora MCP must be connected and authenticated before any
image, video, music, or posting step. If the Unsora MCP tools are missing, stop
before the approval gate and tell the user to connect Unsora MCP first. Do not
silently substitute another generator. `ffmpeg`/`ffprobe` and `python3` (with
`numpy`, `pillow`, `scikit-learn` for QC) must be available locally.

**The style, in one line:** flat matte cardstock worlds — visible cut edges, soft
contact shadows, a strict **three-colour palette and no others** — a tiny figure
dead-centre of a huge symmetrical set, **rendered in 3D, not real stop-motion**,
scenes packed with incident and joined by **paper wipes**, native anime-fight
impact SFX and paper foley, landing on a face.

**What makes it *this* style, not a generic paper cutout:** the look is a
**measurable specification, not a vibe** — palette, exposure and pacing are all
numbers, and every failure is a number drifting off target. State targets as
numbers in the prompt, then **measure the output against them** with the bundled
QC script. Adjectives don't survive the trip to the model; percentages do.

**The failure this skill exists to prevent** is not an ugly clip — it's a
**boring** one: a single slow scene, one camera drift, half the runtime doing
nothing. It looks correct and feels dead. **Density is the deliverable.** Every
shot must carry incident, and the load-bearing device that lets scenes stay dense
without hard-cutting is the paper wipe (see `references/beats.md`).

**The tool split (read this — it is the point of the pipeline):**

- **Everything visual + the music runs on Unsora.** `create_image`
  (`nano-banana-pro`) for hero frames and reference sheets → `create_video`
  (`seedance-2.0`) to animate each frame → `create_music` (`mureka-7.5`) for the
  optional score. Each `create_*` is async; poll its `wait_for_*`.
- **The voiceover never comes from Unsora — it has no TTS.** If the film is
  narrated, the VO comes from **outside**: either the user supplies a VO file, or
  it is synthesized on **whatever TTS tool is connected in the session** (see
  `references/pipeline.md` §VO). The user decides at the front door. Never quietly
  substitute a connector, and never claim Unsora voiced it. A diorama film can
  also run **VO-less** — pure spectacle carried by SFX — which is a perfectly
  valid choice here, unlike a talking explainer.
- **Assembly, wipe-joining, timing, and the mix are ffmpeg**, locally.

**The pipeline, in one line:**

```
elicit params (idea/script · palette · orientation · VO source · music) → read script
  → beat map (one point per shot, flag recurring assets, plan wipe joins) → shot list [+ VO draft]
  → [APPROVAL GATE] →
  [if narrated] VO FIRST: obtain each line's audio → measure real durations → size each shot to its line →
  reference sheets (Unsora create_image, only for assets recurring across shots) →
  per shot: Unsora create_image (3 variants on hero shots) → QC-pick frame → create_video (animate, wipe-bookended) → QC clip →
  [music: Unsora create_music] → ffmpeg concat (cut hidden inside the wipe) → [time VO, duck + mix] → final.mp4
```

If the user wants only prompts for one or a few shots — no generation — run
**Part 1** and output shots in the per-shot format. If they want the film, run
everything.

---

# PART 0 — THE FRONT DOOR (resolve before spending anything)

Infer what you can; ask only what is genuinely missing. On a chat client with
tappable inputs, ask with those; otherwise ask in one short message. Never call a
generation tool before these are fixed and the shot list is approved.

1. **The script / idea.** Paste a script, VO copy, fact list, or beat list
   (`.docx`, `.txt`, `.md`, `.pdf`, fountain, or typed). If it's a bare idea, say
   so and offer to write a beat list first for approval. If the user has no idea,
   propose 3–5 (see `references/beats.md` §Proposing ideas) — **bias hard toward
   action-dense spectacle: myth, folklore, monsters, disasters, transformations.**
2. **Palette.** One of the five tested palettes, locked for the whole film
   (Reference cream/vermilion/navy · Sable lilac/violet/moss · Cobalt
   cream/vermilion/teal · Iris cream/violet/forest · Glacier ice/vermilion/bitumen).
   Show them by name and feel; pick for the user if they'd rather not. This drives
   every prompt after it, so lock it first. To build a sixth, use
   `scripts/palette.py` and read `references/palette.md` first.
3. **Orientation.** Default **`16:9`** — the concentric top-down hero shot reads
   best wide. Offer `9:16` (Shorts/Reels/TikTok) or `1:1`, but warn that the
   overhead ring beat loses power off 16:9.
   **Resolution is ALWAYS `720p` (1280×720 at 16:9) on every clip and holds across
   the whole film — the ONLY exception is the user explicitly stating a different
   resolution (e.g. "make it 1080p").** Do not prompt for resolution, do not raise
   it "for quality", and do not let it vary shot to shot; assume 720p silently and
   pass `resolution: "720p"` on every `create_video` call unless the user overrode
   it. If they do override, use their value on every clip identically.
4. **Voiceover source.** Offer: **(a) narrated, you supply the voiced file**,
   **(b) narrated, voiced on a TTS tool connected in this session** (name only
   ones you actually find), or **(c) VO-less** — spectacle carried by SFX and, if
   wanted, a score. VO-less is common and legitimate for this style. If the user
   named a tool, that's answered. See `references/pipeline.md` §VO.
5. **Music.** Ask plainly whether they want a **background score**. The impact SFX
   are always in (they *are* the action); the musical bed is a taste call. If yes,
   it's a generated **instrumental bed via Unsora `create_music`** (`mureka-7.5`),
   ducked under any VO in the mix. Default to a light anime-action bed for
   spectacle, a restrained underscore for explainer, unless the user says silent.
6. **Length / shape.** Infer from the script. For a bare idea, ask for a target
   runtime or scene count. Multi-minute is the norm for this skill; the VO (if
   any) is the timing spine.
7. **Unsora MCP availability.** Confirm the session exposes `create_image`,
   `wait_for_image`, `create_video`, `wait_for_video`, `create_music`,
   `wait_for_music`. If not, stop and ask the user to connect Unsora MCP.

---

# PART 1 — THE STYLE SYSTEM

## Rule 0 — every shot carries incident, and the paper never lies

Each shot answers *"what does the viewer see happen that they didn't a second
ago?"* — a blow lands, a head erupts, the world tilts overhead into rings, a face
fills the frame. A shot that is "a nice diorama drifting" with no event is filler:
pack it with action, fold it into a neighbour, or cut it. **Test every shot:**
name the event in one clause. If you can't, it's a postcard, not a shot.

Two non-negotiables sit under every shot:

- **The three-colour spec.** Ground / accent / void at fixed luma (197 / 77 / 38),
  roughly 45% / 35% / 20% of the frame as a **time-average across the film, not a
  per-frame lock**. No other hue, no pure white, no pure black. Full maths, the
  five palettes, and how to write it three ways into a prompt: `references/palette.md`.
- **Paper logic.** Sky is stacked bands of card, not gradient; the sun/moon/lamp
  is a flat disc; an explosion is flat shards, never smoke; sparks are punch-out
  confetti; every surface is fully matte, no specular, no gloss. Any new element
  entering a shot must be named as built from the same stacked cardstock in the
  same three colours, or the model reaches for a photoreal object or a fourth hue.
  Full list: `references/beats.md` §Paper logic.

The full visual system (paper logic, palette, exposure swing, camera, motion,
audio, the AI-default negative, and both prompt templates) is in
**`references/style-dna.md`** — read it before writing a single prompt.

## Exposure swings — do NOT lock it

The single most tempting mistake is telling the model to keep exposure and palette
steady. The reference does the opposite: it swings luma ~150 across the film (a
flat, dead clip swings ~56). The palette ratio is a **time-average, not a
per-frame lock** — the film earns that average by swinging hard (void-heavy in one
beat, ground-flooded in another). **Lock the ratio on the hero frame; let the film
swing around it.** Never write "hold the balance every frame" — it bans the swing
and produces exactly the flat clip you were trying to avoid.

## The paper wipe joins everything — including the cuts between clips

The reference films **do not hard-cut and do not cross-dissolve.** The camera
pushes into a blank sheet of cardstock; for ~0.4–0.5s the frame is one flat
colour; then it emerges somewhere new. **Cream wipes into bright scenes, vermilion
wipes into dark ones** — the wipe colour previews the beat it opens.

In a multi-shot film, the wipe does double duty (full grammar in `references/beats.md`):

- **Within a long shot** (a VO line of ~8s+) — one or two internal wipes keep it dense.
- **Between shots — the wipe HIDES the concat cut.** Seedance clips are discrete
  and must **not** be chained. Instead, **bookend the join:** the outgoing clip
  *ends* by rushing into a flat-paper field of the join colour, and the incoming
  clip *opens* on the same flat-paper field before pulling back into its tableau.
  Concatenated, the hard cut lands **between two flat identical-colour frames**, so
  it is invisible — the two half-wipes read as one continuous ~0.5s paper wipe
  across the seam. The join colour is a property of the join, set by the incoming
  beat's brightness, and written into **both** the tail of shot N and the head of
  shot N+1. The **first shot cold-opens** (no leading wipe) and the **last shot**
  lands on its face and holds (no trailing wipe).

## Choosing a subject — find the ring

The most memorable diorama beat is a **top-down where the world resolves into
concentric rings** radiating from the figure at dead centre (Busby Berkeley). It
only works if some scene in the film has a circular geometry to fall into — a
stairwell, a drum, a whirlpool, a bullring, a crater, a vault, a tornado. A
multi-minute film needs **at least one** such hero shot; find the ring hiding in
the story (a crowd becomes a ring around a pit; a duel becomes a ring of
onlookers). Better still, give the figure something to fight *inside* the ring. A
concentric motif you only look at is a postcard; one you can brawl in is a film.

## The two-prompt structure (every shot)

Each shot is two prompts, built from the `references/style-dna.md` templates:

- **Hero-frame prompt** (Unsora `create_image`, `nano-banana-pro`): the shot's
  single strongest frame — its resolved tableau, palette and exposure locked.
  Ordered Subject → Colour budget → Exposure → Medium/paper-logic → Composition →
  Lighting → Constraints. It MUST carry the three-colour spec stated three ways,
  the exposure block, and the AI-default negative (matte stacked cardstock, cut
  edges, no gloss/CGI/smoke/photoreal).
- **Animation prompt** (Unsora `create_video`, `seedance-2.0`): animates that
  frame. Names the frame by role ("the reference image is the opening frame; the
  shot begins on exactly this composition"), restates the palette (the image
  constrains the *first* frame far more than the last), pins every camera move to
  a timecode and forbids pausing, names the wipe bookends with colour + duration,
  and closes with the audio block (native SFX + paper foley, **No music, No
  voice** — VO and score are separate tracks added in the mix).

**Frame-zero rule:** the hero frame is the shot's resolved tableau. Never write
"hold" anywhere — the model obeys literally and freezes on a full-detail frame.
Give it micro-motion instead (drifting dust, a trembling fringe, idling shiver);
something must move every second that isn't a wipe.

---

# PART 2 — THE PRODUCTION PIPELINE

Never skip the approval gate. Read `references/pipeline.md` for exact Unsora
parameters, VO handling, QC, and ffmpeg before generating.

## Phase 1 — Read the script

Accept `.docx`, `.txt`, `.md`, `.pdf`, fountain, or pasted text; read in full,
then classify: **narration script** (prose to be heard — the VO is the timing
spine), **beat/fact list** (one event per line — roughly one shot each), or
**silent spectacle** (a premise to dramatize with no narration). If it's a bare
idea, offer to write a beat list first.

## Phase 2 — Beat map + asset tally + wipe plan

Apply Rule 0. Split into shots where each shot carries **one event**. Give the
film a shape: **cold open on motion** → build of packed scenes → the **concentric
top-down hero shot** → **land on a face**. Tally every recurring **character,
place, and significant object** against its shots — this feeds the reference-sheet
plan (3+ shots → sheet; the film's hero subject → sheet even at 2; recurring place
→ sheet; one-offs → per-shot). For every join, note the **wipe colour** (set by
the incoming beat: cream into bright, vermilion into dark). Assign each shot a
provisional duration (and a VO line, if narrated). Segmentation, the film arc, the
wipe grammar, density, subject choice, and the VO-timing table are in
`references/beats.md`.

## Phase 3 — Shot list [+ VO draft], then stop

Present the shot list as a table and **wait for a yes**:

| # | Event / beat | Setting & palette use | Wipe in (colour) | Sec | Assets | VO line (if narrated) |

Below the table: the **VO script** (all lines, read start to finish — only if
narrated), the **asset/sheet plan** (name, kind, shots, one sentence of locked
design), the **locked palette** (the three hexes + which colour opposes the
accent), then total runtime, shot count, sheet count, orientation, resolution
(720p unless asked), VO source, and — plainly — that generating spends **Unsora**
credits for images/clips/music (plus whatever the VO tool costs), and takes
roughly `clips × ~7 min` (Seedance) plus a bit per sheet, per VO line, and for the
music. Do not call a single generation tool before the user says go. If the film
is narrated, treat the VO as public copy: vary sentence length, cut filler, and
read it aloud so it doesn't land as generic AI narration (this skill is
self-contained — do this inline, don't reach for another skill).

## Phase 4 — Voiceover first (the timing spine, if narrated)

Skip entirely for a silent film. Otherwise, per the source chosen at the front
door (mechanics in `references/pipeline.md` §VO):

- **Supplied file** → ask for it now; split per beat, or read the whole file's
  duration and lay lines against the beats.
- **Session TTS tool** → find the TTS tool actually available, audition/pick ONE
  voice, synthesize **one call per VO line**, hold that voice across every line,
  download each to `vo/vo_NN.<ext>`.
- **None available and no file** → say so before the gate; offer the silent cut or
  a picture-lock the user narrates themselves.

Then `ffprobe` each line's real duration and **set each shot's `duration` to fit
its line** (word budget ≈ 2.4 words/sec; ~0.3s air each end; Seedance floor 4s).
For a silent film, set durations from the action instead — long enough to land the
event plus its wipe, packed with micro-motion. This measured `(shot → seconds)`
table is what Phase 6 generates against.

## Phase 5 — Write the prompts

For each approved shot, write both prompts with the `style-dna.md` templates and
show them in the per-shot format below before generating. Shots that use a sheet
get the sheet-binding lines and describe only pose/arrangement for that element.
State the palette three ways and the exposure block in **every** hero-frame
prompt; restate the palette in **every** animation prompt. Cheap to fix on the
page, expensive after. Run the `style-dna.md` checklist per shot.

**Per-shot output format:**

1. `### SHOT N — [TITLE]` + one line: the event and how this shot lands it. If it
   uses sheets, a second line: `Assets: [name → role]`.
2. **Breakdown** — duration & pacing; subject(s); action & blocking; setting &
   palette use; wipe in/out (colour + duration); composition; camera & timecoded
   moves; exposure; the audio block; THE HERO FRAME.
3. **Hero-frame prompt** — one fenced code block.
4. **Animation prompt** — one fenced code block.

## Phase 6 — Generate + QC (Unsora)

Read `references/pipeline.md` first. Reference sheets once each (before any shot),
QC and show them to the user. Then per shot:

- `create_image` → `wait_for_image`. On the **hero shots** (the establish, the
  concentric top-down, the final face), generate **3 variants** and pick on
  numbers: `python scripts/qc.py frame v1.png v2.png v3.png --ground '#…' --accent
  '#…' --void '#…'`. On routine shots, one frame + QC, re-roll only on a fail. A
  bad plate costs ~1 min to re-roll and is far cheaper than animating it.
- `create_video` (hero URL as `image`) → `wait_for_video`. `aspectRatio`
  **identical** in both calls (a mismatch crops the frame). `resolution: "720p"`
  on every call unless the user asked for 1080p; hold it across all shots.
  `generateAudio: true`. Pass a shot's sheet URLs in `referenceImages` of **both**
  calls (slot math: in the video call `image` = @Image1, first reference =
  @Image2). Do **not** chain last-frame → next-first-frame — the wipe bookends are
  the join.
- QC the clip: `python scripts/qc.py video clips/shot_NN.mp4 --ground … --accent …
  --void …`. Report the numbers honestly, including failures (exposure swing, wipe
  count, drone, palette drift). Verify the shot opens/closes on a flat field if it
  bookends a join.

Keep the manifest updated so a failed shot regenerates alone. Never re-roll a
`done` shot "for consistency."

## Phase 7 — Music (optional, Unsora)

If a bed was requested: `create_music(model="mureka-7.5", prompt=<named
instruments, tempo, a build to a hit — e.g. "taiko drums and shakuhachi building
to a hard stab", or a restrained editorial underscore for explainer; instrumental,
no vocals>)` → `wait_for_music` → download. Vague "epic music" comes back as drone;
name the instruments and the arc. The bed rarely matches runtime — loop or trim it
in the mix.

## Phase 8 — Assemble, join, [time VO], mix, deliver

ffmpeg. Concat clips in manifest order (zero-pad shot numbers) — because each join
is wipe-bookended, the cut lands inside a flat field and needs no crossfade.
**Verify wipe continuity** at each seam (the join colour matches on both sides). If
narrated, lay each `vo_NN` at its shot's start offset over the clips' native SFX
(kept low so impacts read), with the **music bed ducked under the VO** (sidechain);
for a silent film the SFX carry, and the score (if any) sits under them. Verify
final duration against the picture cut — a dropped clip yields a valid file that is
simply short. Exact ffmpeg (concat, VO timing + duck/mix, brand-bug overlay,
verification, resume) is in `references/pipeline.md`. Present `final.mp4` with the
host's native file delivery. Offer, but never perform unprompted, any social post
(Unsora `create_post` — needs its own explicit yes).

## Failure handling (universal; `style-dna.md` adds style-specific rows)

| Symptom | Cause | Fix |
|---|---|---|
| Clip is boring — one slow scene, half the runtime dead | Not enough scenes/wipes; "hold" in prompt | Pack more events; name wipes explicitly with colour + duration; purge every "hold", give micro-motion |
| Look turns glossy / photoreal / CGI, or grows smoke & volumetrics | The diorama AI-default leaked | Add the medium sentence + AI-default negative (stacked matte cardstock, visible cut edges, flat shards not smoke, no gloss/specular/lens-flare) to every prompt |
| Frame or clip drifts dark (esp. interiors, close-ups) | No exposure block, or a global exposure lock | Put the exposure block in from attempt one; pin the close-up "as bright as the opening"; never write "hold the balance every frame" |
| Palette drifts / a fourth colour appears / accent floods | Ground collapsed; too many accent cues; void doesn't oppose accent | State palette three ways; name the removed colour ("NO BLUE — the dark here is green"); check one colour opposes the accent ~140°; keep accent to one source |
| Camera move never happens / arrives late | Move unpinned | Pin every crane/push to a timecode and forbid pausing ("the crane starts at 0:09 and does not pause or slow") |
| A recurring face/place looks different between shots | No sheet, or sheet not passed/bound | Generate the sheet once; pass its URL in `referenceImages` of both calls, bind by slot in both prompts |
| A low continuous drone appears under the biggest move | Seedance adds it unasked | Ban the SUSTAIN not the frequency: "no sustained tone, no cinematic drone, no continuous hum"; keep "heavy card dropped flat" for weight |
| Seedance invents a voice | VO leaked into the video prompt | Keep VO out of every video prompt; `No voice.` is load-bearing — the VO is a separate track |
| Visible hard cut at a join | Bookends missing or wipe colours mismatched | Re-render so the outgoing clip ends and the incoming opens on the SAME flat-paper colour; verify head/tail flat fields |
| "Unsora voiced it" | Unsora has no TTS | State the real source (session TTS tool, or supplied file); never attribute the VO to Unsora |
| Assembled film is short | A clip failed and got skipped | Check the manifest, regenerate the missing shot |
| A clip is the wrong size / concat won't stream-copy | Resolution drifted off 720p | Every clip must be `720p` (1280×720) unless the user overrode it; regenerate the odd clip with `resolution: "720p"` and hold it across all shots |

## Reference files

- `references/style-dna.md` — the full diorama visual system: paper logic, the
  three-colour spec, the exposure swing, camera, limited motion, the audio SFX
  vocabulary, the AI-default betrayal + its negative, both prompt templates, the
  reference-sheet formats, the wipe-bookend wording, style-specific failure rows,
  and the per-shot checklist. **Read before writing any prompt.**
- `references/palette.md` — the palette maths (luma geometry, why the accent is
  nearly locked to red), the five tested palettes + one documented failure, how to
  state a palette three ways in a prompt, and how to build a sixth with
  `scripts/palette.py`.
- `references/beats.md` — the paper-wipe device (within-shot and the between-shot
  bookend join), the film arc, density, cold-open, subject/ring choice, proposing
  ideas, camera, VO-first timing, and a worked script → shot map.
- `references/pipeline.md` — exact Unsora parameters and polling (sheets, hero
  frames + variant QC, clips, `create_music`, reference-slot ordering); the VO
  paths (supplied file / session TTS); the manifest schema; the QC steps; every
  ffmpeg command (wipe-join concat, VO timing + duck/mix, brand bug, verification,
  resume-after-failure); and the optional Unsora posting flow.
- `scripts/qc.py` — measures hero frames (`frame`) and finished clips (`video`)
  against the palette/exposure/wipe/audio targets. `python scripts/qc.py --help`.
- `scripts/palette.py` — solves for a new three-colour palette at the right luma
  values. `python scripts/palette.py --ground-hue … --accent-hue … --void-hue …`.
