diff --git a/README.md b/README.md index 4cae16b1..09abbd99 100644 --- a/README.md +++ b/README.md @@ -10,21 +10,21 @@ HocusPocus exists for creators who do not want a one-off prompt machine. It is a 1. **Studio (sidebar)** — choose an image, video or audio model; write the prompt; add references/LoRAs; then generate. Use it when you want direct, manual control over one asset. The output appears in the gallery and is reusable everywhere else. 2. **Director (sidebar)** — select Music Video, Short Film, Trailer or a story-driven workflow and describe the outcome. The LLM turns the brief into reviewable shots, prompts and references. Choose manual review for control or automatic mode for a complete recoverable pipeline. -3. **Gallery: All, Images, Videos, Audio, Videoclips, Trailers and Chapters** — browse results by kind. Open an item to inspect it; use it as a reference, send it to an editor, or keep it in the active workspace. +3. **Gallery: All, Images, Videos, Audio, Videoclips, Trailers and Chapters** — browse results by kind. Open an item to inspect it; use it as a reference, send it to an editor, or keep it in the active output folder. 4. **3D** — create a mesh from text, one image or the four front/left/right/back views. You can upload references or select existing HocusPocus images. Export GLB for later animation or 3D-video composition. -5. **3D Video** — place GLBs, images and effects in a controllable scene and render an MP4/WebM composition. Use it for camera moves that must be exact rather than invented by a video model. +5. **3D Video** — place GLBs, images, Character Kits and effects in a controllable scene and render an MP4/WebM composition. Use it for camera moves that must be exact rather than invented by a video model, and for 2D cutout dialogue through Face Rig mouth overlays. 6. **Animate** — rig a compatible static GLB and apply procedural or AI-assisted motion. Export the rigged model or bring it back to 3D Video. -7. **Character Creator** — upload one subject image. H3 makes a centered 360° turntable; select/re-take the best front, left, back and right frames, then build a Hunyuan multi-view mesh. It is the quickest route from character art to a 3D asset. +7. **Character Creator** — upload one subject image. H3 makes a centered 360° turntable; select/re-take the best front, left, back and right frames, then build a Hunyuan multi-view mesh. **Create / open CharacterKit Face Rig** hands a saved view to **3D Video** as a 2D puppet (not a mesh). It is the quickest route from character art to a 3D asset or a reusable cutout; see the [Character Kits / Face Rig guide](docs/character-kits/HOWUSEIT.md). 8. **Story Lab** — write the production bible: premise, world rules, locations, cast, relationships and beats. Review each field, then hand the approved canon to Comic Studio, Director, trailers or music-video production. 9. **Series Lab** — turn approved canon into seasons, episodes, scenes and shots. Use it when continuity has to survive several episodes and you need attempts and references tracked per shot. -10. **Workspaces** — separate client work, worlds and experiments. The workspace selector changes the active output folder; this tab is also the place to inspect and resume Director production threads. +10. **Output folders and Workspace collections** — the output-folder selector changes the active physical save directory. Explicit Workspace collections group related projects, assets and Productions without becoming a directory; the Workspaces tab inspects and resumes Director production threads in the active output folder. 11. **Hoja de estilos** — save visual rules, palettes and prompt language. Apply the sheet so images, comics and shots share a recognizable art direction. 12. **Comics** — plan pages and panels, revise dialogue, generate/re-generate panel art, then export PDF, CBZ or PNG. Use **Comic → AI film** to turn approved panels into a film plan without losing the comic canon. 13. **Video Editor** — import generated or uploaded clips, trim, split, reorder and export an MP4. It is the finishing room for material that is already good and should not be regenerated. 14. **Edits and Multi-clip** — use Edits for targeted retakes, outpaint and prompt-driven changes; use Multi-clip for longer sequences with prompt-by-prompt control and overlapping continuity. 15. **Footer, Productions and Settings** — the footer shows active jobs, their ETA and history; Productions preserves Director pipelines and recovery actions. The gear controls services, models, performance, themes and explicit mature/experimental options. -For a typical project: create the canon in **Story Lab** (or begin in **Studio**), approve a character or location, produce a few clips in **Director**, refine them in **Video Editor**, and keep all source assets in the same **Workspace**. Every step can also start from an existing image, video, audio file or GLB. +For a typical project: create the canon in **Story Lab** (or begin in **Studio**), approve a character or location, produce a few clips in **Director**, refine them in **Video Editor**, and keep source assets in the same **output folder**, optionally linking the work to a logical **Workspace collection**. Every step can also start from an existing image, video, audio file or GLB. ## What it does @@ -138,7 +138,7 @@ For a screenshot-led, end-to-end walkthrough, see **[HocusPocus / Experimental: - **Guided mode** creates reviewable field-level drafts and locks production until the relevant story, world, cast identities, relationships, and structure are approved. **Automatic mode** runs the same checkpointed pipeline, then offers first-look world, location, and character concepts. - Choose HocusPocus internal, DeepSeek V4 Pro/Flash, MiniMax M3/M2.7/M2.7 Highspeed, OpenAI, or a custom compatible writing agent inside the story itself. Concept-art generation has its own independent HocusPocus/MiniMax selector. - Character cards combine role, desire, need, flaw, arc, dialogue voice, wardrobe, visual invariants, negative prompts, multiple references, and a selected primary identity image. -- Export/import a `.storypack` with the editable JSON and available visual assets. Each workspace has a multi-story autosaved library; generated plans and local concept jobs can resume from durable checkpoints after interruption. +- Export/import a `.storypack` with the editable JSON and available visual assets. Each output folder has a multi-story autosaved library; generated plans and local concept jobs can resume from durable checkpoints after interruption. A logical Workspace collection can link those projects without owning their files. - **Productions** offers a review-first hand-off and a complete one-click generation for both media. Comic opens **Director → Comic** and creates a self-contained chapter rather than retelling the master plot; four pages remain the quick-test default, while page count and panels per page are configurable up to the Director limits. Short Film opens **Director → Short Film → Story** with an editable target duration and independently selectable image and video models, and inherits the Story project's selected writing provider instead of silently falling back to the global LLM. Its shot frames can use a local HocusPocus image model or the external MiniMax Image-01 API; the latter does not consume local VRAM and is distinct from the local MiniMax H3 video runtime. Both productions receive the full editable canon, structured cast, locations and labelled visual references. Character images remain attached through planning and MiniMax `image-01` uses the visually prioritised character as its single supported identity reference per request. Adaptation history preserves the selected models when reopening the staged target, or can restore its exact source as a new editable copy. - New **Videoclip** songs default to local **ACE-Step 1.5 XL** (`ace_step_v1_5_xl_sft_lm_4b`). MiniMax `music-2.6` / `music-3.0` stay available on the song model selector. - **Tráiler cinematográfico** is a standalone Story Lab project type beside **Videoclip**, so movie trailers never require a song. Its four-stage planner creates the concept, protagonists, world and a 6–12-beat trailer arc, then opens the dedicated 15–180 second Trailer Creator. It exposes theatrical, teaser and character formats; narration, dialogue or visual-only storytelling; spoiler and intensity controls; optional minimal title cards; and an editable six-part timed arc from cold open to unresolved final hook. Visual generation can create start frames, route approved references directly through H3 Ref2VA, or run as pure text-to-video without generating or sending any image. A trailer can be reviewed in Director or generated as a recoverable ordered pipeline, then replayed, regenerated clip-by-clip and joined from Story Lab's Assembly view. @@ -153,7 +153,7 @@ Build a comic as an editable production rather than a single flattened generatio - Comic and Video Editor drafts autosave locally; saved comics remain backward-compatible with older version-2 project JSON. ### ✂️ Video Editor — cut Lab clips without regenerating -The **Video Editor** tab assembles existing workspace or uploaded clips (H3 MP4s, compositor WebM, Series handoffs, comic animatics). Timeline trim, split, reorder, play-from-selection, and playhead scrub are local; only probe, thumbnails, frame grabs, and FFmpeg export hit the server. +The **Video Editor** tab assembles existing output-folder or uploaded clips (H3 MP4s, compositor WebM, Series handoffs, comic animatics). Timeline trim, split, reorder, play-from-selection, and playhead scrub are local; only probe, thumbnails, frame grabs, and FFmpeg export hit the server. - Import from the gallery card, **From HocusPocus** picker, drag-and-drop (`POST /api/v1/upload`, 500 MB max), or a Series/comic handoff. - Export is a queued FFmpeg job (`POST /api/v1/video-editor/export`, 1–100 clips, even 240–3840 resolution, fps 24/25/30/50/60). Cancel waits for the current FFmpeg subprocess. @@ -191,7 +191,7 @@ Three themes, switchable in Settings → System: ### 🧊 Native Hunyuan3D Studio HocusPocus includes an integrated **3D** section for text-to-3D, image-to-3D, and multi-image reconstruction. Hunyuan runs in an isolated environment inside HocusPocus, so its older Diffusers stack cannot conflict with the audio/video models. Each worker exits after export and releases its CUDA context and VRAM. -Reference slots accept both images uploaded from disk and existing images selected directly from the active HocusPocus workspace. This applies to the single front reference, all four multi-view slots, and the texture reference used by **Retexture GLB**. +Reference slots accept both images uploaded from disk and existing images selected directly from the active HocusPocus output folder. This applies to the single front reference, all four multi-view slots, and the texture reference used by **Retexture GLB**. Included geometry variants: @@ -284,10 +284,10 @@ job = requests.post(f"{base}/api/v1/model3d/generate", json={ status = requests.get(f"{base}/api/v1/model3d/status/{job['job_id']}").json() ``` -### 📂 Workspaces (output directories) -Multiple isolated output directories with a quick switcher in the sidebar. Useful for separating client projects, NSFW vs SFW, or experiments. Pinned and favorited outputs are tracked per workspace. +### 📂 Output folders and Workspace collections +Multiple isolated physical output folders with a quick switcher in the sidebar are useful for separating client projects, NSFW vs SFW, or experiments. Pinned and favorited outputs are tracked per output folder. A separate Workspace collection is an explicit logical grouping of project, asset and Production IDs; it does not create, move or delete files. See the [domain model and asset provenance contract](docs/development/DOMAIN_MODEL_AND_ASSET_PROVENANCE.md). -The gallery **Workspaces** tab is a different feature: a Director **generation-thread** dashboard for the *active* output directory (inspect the shot queue, rewrite selected prompts from one instruction, toggle per-shot soundtrack drive vs mute, resume, rejoin). Operator contract: **[Workspaces tab](docs/workspaces/HOWUSEIT.md)**. +The gallery **Workspaces** tab is a different feature: a Director **generation-thread** dashboard for the *active* output folder (inspect the shot queue, rewrite selected prompts from one instruction, toggle per-shot soundtrack drive vs mute, resume, rejoin). Operator contract: **[Workspaces tab](docs/workspaces/HOWUSEIT.md)**. ### 🔒 Mature mode + experimental gate - **NSFW mode** is opt-in with a disclaimer step. Disabled by default. Gates uncensored model variants, NSFW LoRAs in the CivitAI browser, and the Settings → Services NSFW toggle. @@ -305,7 +305,7 @@ The version you are running is shown next to the HocusPocus title in the UI. To **First HocusPocus preview** - HocusPocus now has its own product versioning, separate from the upstream Maestro lineage. - The local creation studio introduces its final transparent scribe icon, launch intro, and English-first identity. -- Image, video, sound, 3D, comics, games, and story tools remain available in the same local workspace. +- Image, video, sound, 3D, comics, games, and story tools remain available in the same local output folder. ### Pre-HocusPocus lineage @@ -439,7 +439,7 @@ After clicking **Start**, the launcher shows an **Open Web UI** button once the - **Activity footer** — persistent live job progress and access to current or past **Productions** - **Settings drawer** (gear icon) — model visibility, performance auto-tune, services (LLM, API keys, NSFW, theme) - **Pinokio menu** — Update, Reset, Install Inpaint Support, LoRA folder shortcuts -- **Operator guides** — [HOWUSEIT index](docs/HOWUSEIT.md) (Video Editor, Workspaces tab, 3D compositor) +- **Operator guides** — [HOWUSEIT index](docs/HOWUSEIT.md) (Video Editor, Workspaces tab, 3D compositor, Character Kits) ## Sharing on the local network diff --git a/app/docs/API.md b/app/docs/API.md index 3be63a42..b8682236 100644 --- a/app/docs/API.md +++ b/app/docs/API.md @@ -603,7 +603,7 @@ does not apply local image LoRAs. The credential is read only from Settings → Services; it is not accepted in pipeline payloads or persisted in output metadata. -Comic recovery checkpoints are workspace-scoped and durable. Create one with +Comic recovery checkpoints are output-folder-scoped and durable. Create one with `POST /api/v1/comics/history`, list versions with `GET /api/v1/comics/history` (optionally `?comic_id=...`), and load a version with `GET /api/v1/comics/history/{snapshot_id}`. Identical consecutive @@ -720,13 +720,72 @@ curl -X POST "$MAESTRO_URL/api/v1/characters/describe-refs" \ `kind` is `character` or `object` (default `character`). `roles` are `subject`, `face`, `outfit`, `extra`, or `accessory`; missing/invalid roles become `subject` for index 0 and `extra` otherwise. Response: `{ "a_prompt", "kind" }`. `400` if `image_paths` is empty, a file is missing, or the API key is unset. +## Character Kits + +Character Kits are reusable 2D cutout puppets. Their current HTTP routes are +scoped to a physical output folder. The query/body field is named `workspace` +for compatibility; it is **not** the ID of a logical Workspace collection. +Logical collections use `/api/v1/workspace-collections` and only group project, +asset, and Production IDs. See [`docs/character-kits/HOWUSEIT.md`](../../docs/character-kits/HOWUSEIT.md) +for the operator workflow and [the domain contract](../../docs/development/DOMAIN_MODEL_AND_ASSET_PROVENANCE.md). + +Set the base URL in the examples to your running HocusPocus instance: + +```bash +export HOCUSPOCUS_URL=http://127.0.0.1:7860 +``` + +- `GET /api/v1/character-kits/library?workspace=default` — reads the normalized + `{output-folder}/.character-kit-library-v1.json`, or returns an empty + `{ "version": 1, "revision": 0, "activeId": "", "kits": {} }` when it is + missing. +- `PATCH /api/v1/character-kits/library/kits/{kit_id}` — creates or replaces one + kit with `{ workspace, baseRevision, kit, makeActive? }`; neighbours are not + replaced. `makeActive` defaults to true. A stale revision returns `409` with + `code: character_kit_revision_conflict`, `expectedRevision`, and + `currentRevision`. +- `DELETE /api/v1/character-kits/library/kits/{kit_id}` — body + `{ workspace, baseRevision }`; `404` if the kit is absent. Source files are + intentionally retained. +- `POST /api/v1/character-kits/face-rig/cleanup` — body + `{ workspace, source, padding? }`, where `padding` is 0–64 (default 8). The + endpoint runs rembg U2Net + crop-to-alpha, writes a new PNG, and never + overwrites `source`. Sources must be inside uploads or the selected output + folder; disallowed/missing images return `400`/`404`. + +The output-folder token is `default` or `[A-Za-z0-9][A-Za-z0-9_-]*`. Kit mouth +keys are `closed`, `small`, `wide`, and `round`; eye keys are `open` and +`blink`. `blob:` sources are rejected, and the UI-only `lookNotes` field is +stripped when the kit is normalized for persistence. + +```bash +curl "$HOCUSPOCUS_URL/api/v1/character-kits/library?workspace=default" + +curl -X PATCH "$HOCUSPOCUS_URL/api/v1/character-kits/library/kits/luma" \ + -H "Content-Type: application/json" \ + -d '{ + "workspace": "default", + "baseRevision": 0, + "kit": { + "version": 1, "id": "luma", "name": "Luma", "style": "cutout", + "base": { + "id": "luma-base", "name": "Luma base", "source": "luma-base.png", + "kind": "image", "alphaStatus": "transparent", "reviewState": "approved" + }, + "poses": {}, "mouth": {}, "eyes": {}, + "anchors": { "base": { "mouth": { "offsetX": 0, "offsetY": -18, "scale": 0.05, "rotation": 0 } } }, + "provenance": [] + } + }' +``` + ## Director pipeline threads -These routes always use the server active workspace. They do not accept `?workspace=`. +These routes always use the server active output folder. They do not accept `?workspace=`. -- `GET /api/v1/director/pipelines` / `GET /api/v1/director/pipelines/active` / `GET /api/v1/director/pipelines/{pid}` — list or load. `{pid}` hydrates an empty `clips` array from `clip_plans` or `planned_clips` and sets `queue_source` to `clips`, `clip_plans`, or `planned`. +- `GET /api/v1/director/pipelines` / `GET /api/v1/director/pipelines/active` / `GET /api/v1/director/pipelines/{pid}` — list or load. `{pid}` hydrates an empty `clips` array from `clip_plans` or `planned_clips` and sets `queue_source` to `clips`, `clip_plans`, or `planned`. List accepts `limit` and `offset` (newest first); `limit=0` (the default) returns the full list, while the Workspaces tab pages 8 at a time and uses `total` for “load more”. - `PUT /api/v1/director/pipelines/{pid}/clips/{clip_index}/prompt` — optional `video_prompt`, `image_prompt`, `soundtrack_drive`. `true` writes an `audio_driven` / `lip_sync_critical` plan; `false` writes `music_driven` and clears `_director_dialogue_beats`. `409` while the pipeline is active. - `POST /api/v1/director/pipeline/{pid}/resume` and `POST /api/v1/director/pipeline/{pid}/continue` use the singular `pipeline` path. - Batch prompt rewrite is UI-only: loop `POST /api/v1/llm/generate` (local LLM) then PUT the chosen prompts. -Operator notes: `docs/video-editor/HOWUSEIT.md` and `docs/workspaces/HOWUSEIT.md`. +Operator notes: `docs/video-editor/HOWUSEIT.md`, `docs/workspaces/HOWUSEIT.md`, and `docs/character-kits/HOWUSEIT.md`. diff --git a/docs/3d-video-compositor/HOWUSEIT.md b/docs/3d-video-compositor/HOWUSEIT.md index 10ba3ce9..7fc7e311 100644 --- a/docs/3d-video-compositor/HOWUSEIT.md +++ b/docs/3d-video-compositor/HOWUSEIT.md @@ -1,6 +1,6 @@ # HOWUSEIT — 3D Video compositor (Scene Animator) -Agent operations guide for Loreframe Lab’s programmatic compositor. +Agent operations guide for HocusPocus’s programmatic compositor. This document is the agent operations guide. The **3D Video** tab provides the Recipe runner (intent → JSON → assets → editable scene → MP4), template mounting and the selected-layer copilot. Story Lab / trailers / videoclips remain separate consumers. @@ -19,8 +19,9 @@ A **layered compositor**, not MiniMax H3. | Scene Animator | **3D Video** | Stack images, videos, GLBs, camera, rain/fog/etc. Animate them. Record a clip | | MiniMax H3 | Studio / Story Lab | Native video + stereo audio (acting, dialogue, locations) | | Video Editor | **Video Editor** | Join compositor clips with H3 clips | +| Character Kits | **3D Video** sidebar | Reusable 2D cutout puppets + Face Rig mouth overlays. Operator guide: [Character Kits](../character-kits/HOWUSEIT.md) | -Use the compositor when you need **controllable motion of a known object** over plates: a ship crossing stars, a UFO rising behind mountains, a logo flying in, rain over a still. Use H3 when you need **performance, speech, or a living location**. Mix them: H3 for people/places, compositor for the vehicle insert, Video Editor to cut them together. +Use the compositor when you need **controllable motion of a known object** over plates: a ship crossing stars, a UFO rising behind mountains, a logo flying in, rain over a still. Use H3 when you need **performance, speech, or a living location**. Mix them: H3 for people/places, compositor for the vehicle insert, Video Editor to cut them together. Use a **Character Kit** when the known object is a graphic puppet that must speak with mouth overlays—not a Hunyuan mesh and not H3 lip-sync. Do **not** ask H3 to “keep this exact GLB flying on a perfect path.” H3 will invent a new ship. The compositor keeps the mesh. @@ -86,7 +87,7 @@ Camera shake lives on the **camera** layer: `animation.shake = { amount, frequen 4. Set scene size (match the H3 aspect if you will cut together: 16:9 e.g. 1280×720 or 960×544). 5. Assign motion presets (see §7). 6. Optional: camera preset + atmosphere. -7. **Save to Loreframe Lab** → writes `*.scene.json` + preview PNG (gallery tab **Scenes**). +7. **Save to HocusPocus** → writes `*.scene.json` + preview PNG (gallery tab **Scenes**). 8. **Export MP4** → validated H.264 MP4 in **Videos**. Then import it into **Video Editor** with H3 clips. Motion JSON can be imported separately (2 MB max) via the panel’s movement loader. @@ -95,7 +96,10 @@ Motion JSON can be imported separately (2 MB max) via the panel’s movement loa ## 5. Asset APIs (phase 1, fully scriptable) -Base URL: the running Lab (`http://127.0.0.1:`). Workspace query: `?workspace=default` on file URLs; **never** put that query into probe/source filenames. +Base URL: the running HocusPocus instance (`http://127.0.0.1:`). The +legacy `workspace` query selects a physical output folder on file URLs; +**never** put that query into probe/source filenames. It is not a logical +Workspace collection selector. ### 5.1 List outputs @@ -184,9 +188,10 @@ Direct Hunyuan API: } ``` -Image values are workspace filenames, upload names, or `/api/v1/file/...` **without** `?workspace=` (the resolver strips it, but filenames are safer). +Image values are output-folder filenames, upload names, or `/api/v1/file/...` +**without** `?workspace=` (the resolver strips it, but filenames are safer). -In **Hunyuan3D Studio**, every reference slot offers both **Upload** (local disk) and **Loreframe** (images already stored in the active workspace). The selected Loreframe filename is sent with that workspace, so no duplicate upload is needed. +In **Hunyuan3D Studio**, every reference slot offers both **Upload** (local disk) and **HocusPocus** (images already stored in the active output folder). The selected HocusPocus filename is sent with that output-folder token, so no duplicate upload is needed. Poll `GET /api/v1/model3d/status/{job_id}`. Result `filename` is a `.glb`. @@ -228,10 +233,25 @@ Without a real PNG preview the endpoint rejects the body. ### 5.8 Video Editor (join plates) -Import H3 MP4s + recorded WebM. Multi-select in **From Loreframe Lab**. Export MP4. +Import H3 MP4s + recorded WebM. Multi-select in **From HocusPocus**. Export MP4. Director / Series **auto-joins** (not the editor) use a 0.5 s last-frame freeze + ~0.4 s crossfade when there is no driving audio. Full probe/export/`result_kind` contract: [Video Editor / mixes](../video-editor/HOWUSEIT.md). +### 5.9 Character Kits (2D puppets) + +Character Kit library: `GET`, `PATCH`, and `DELETE` +`/api/v1/character-kits/library…`. Face Rig overlay cleanup: +`POST /api/v1/character-kits/face-rig/cleanup`. Both routes use the physical +output-folder token in their `workspace` field; it is not a logical Workspace +collection ID. Character Creator can hand a saved view into this editor via +`hocuspocus:character-kit-face-rig-handoff`, but it does not generate visemes. + +Only **approved** poses and overlays mount into the scene or enter Recipe +inventory (`APPROVED_CHARACTER_KIT`). Spoken cutout dialogue is persisted as +`scene.dialogueBeats` and compiled into held/snap opacity keyframes; it is not +phoneme-perfect lip-sync. Read the full CAS, review, mouth-pack, and dialogue +contract in [Character Kits / Face Rig](../character-kits/HOWUSEIT.md). + --- ## 6. Layer cookbook (what to stack) @@ -451,11 +471,12 @@ Need both in one sequence? Do this **in the browser tab**, not as a Python overnight script. UI: `SceneRecipePanel`. Code: `ui/src/lib/sceneRecipe.ts`, `ui/src/lib/sceneRecipeAssets.ts`. 1. **Interpretation contract**: the selected LLM receives a closed JSON Schema plus a multilingual virtual-production guide. It silently separates subjects, setting, chronological beats, format, camera and atmosphere; generation prompts are written in concise cinematic English while proper names and quoted dialogue are preserved. -2. **Validation and repair**: local llama-server output is grammar-constrained. Other providers receive the exact schema in context. Loreframe then validates unique ids, asset/layer compatibility, supported presets, rig clips and references; one malformed response gets a bounded correction pass before any GPU job starts. +2. **Validation and repair**: local llama-server output is grammar-constrained. Other providers receive the exact schema in context. HocusPocus then validates unique ids, asset/layer compatibility, supported presets, rig clips and references; one malformed response gets a bounded correction pass before any GPU job starts. 3. **Manual**: pick GLBs/plates already in Outputs, **Write recipe**, **Compose**. Output sidecar prompts and embedded clip names are passed to the LLM as untrusted inventory descriptions, so it can understand assets whose filenames are vague. A requested rig profile is applied even to a manually loaded static GLB. Edit the inspector, then Record. 4. **Auto**: **Generate + compose** creates missing plates/meshes. One `identity` per object — a UFO series uses **one** GLB and several `shots[]`. Static environments use image plates; inherently moving scenery can use an H3 video plate. Rain, fog, snow and particles use procedural effects instead of redundant generated overlays. Default `record`/`save` are false so you preview first. 5. GPU jobs and Hunyuan poll with timeouts; **Cancel** aborts the run. If Lab dies (segfault), the runner errors instead of spinning forever. 6. After compose, switch shots in the recipe panel without regenerating the mesh. A recipe rig `clip` is mounted into the Scene Animator and disables unintended turntable spin. +7. Approved Character Kits enter inventory as `APPROVED_CHARACTER_KIT` rows. Spoken cutout shots need a `speech` audio entry and a top-level `dialogueBeats` row whose `mouthLayerIds` name the overlays. Keep body and face pieces from the same kit ID. Keep the 3D Video tab visible while it records. Browser capture is validated and published as H.264 MP4; import that output in Video Editor to join it with @@ -479,6 +500,7 @@ H3 clips. | `app/services/rig_service.py` | Rig profiles and clips | | `docs/minimax-h3-prompting.md` | H3 prompt dialect | | `docs/video-editor/HOWUSEIT.md` | Cut compositor WebM with H3 MP4s; mix kinds | +| `docs/character-kits/HOWUSEIT.md` | Character Kit library, Face Rig, and cutout dialogue | | `ui/src/features/characters/orbitPrompt.ts` | Orbit A/B prompts and still-frame indices | When in doubt: **one identity per mesh, one path per compositor shot, H3 never draws that mesh.** diff --git a/docs/HOWUSEIT.md b/docs/HOWUSEIT.md index b5bf8f4e..9ad12ef3 100644 --- a/docs/HOWUSEIT.md +++ b/docs/HOWUSEIT.md @@ -1,7 +1,7 @@ # HOWUSEIT -Operator guides for Loreframe Lab subsystems. Prefer these over inventing a second workflow. +Operator guides for HocusPocus subsystems. Prefer these over inventing a second workflow. - **Video Editor** (import Lab clips, timeline, export) and **assembled mixes** (`result_kind` gallery tabs): [`docs/video-editor/HOWUSEIT.md`](video-editor/HOWUSEIT.md) -- **Workspaces tab** (Director generation threads — not the output-directory switcher): [`docs/workspaces/HOWUSEIT.md`](workspaces/HOWUSEIT.md) +- **Workspaces tab** (Director generation threads — not the output-folder selector): [`docs/workspaces/HOWUSEIT.md`](workspaces/HOWUSEIT.md) - **3D Video compositor** (Hunyuan meshes + plates + camera + rain/fog, mixed with MiniMax H3): [`docs/3d-video-compositor/HOWUSEIT.md`](3d-video-compositor/HOWUSEIT.md) diff --git a/docs/character-kits/HOWUSEIT.md b/docs/character-kits/HOWUSEIT.md new file mode 100644 index 00000000..9c635e9d --- /dev/null +++ b/docs/character-kits/HOWUSEIT.md @@ -0,0 +1,368 @@ +# HOWUSEIT — Character Kits and Face Rig + +Operator guide for reusable **2D cutout puppets**: one reviewed body or pose, +mouth overlays, an optional blink, and pose-local face anchors. This is not +Character Creator's Hunyuan turntable and not MiniMax H3 lip-sync. + +UI: **3D Video** sidebar (`SceneAnimatorPanel` → Character Kits). Code: +`ui/src/lib/characterKit.ts`, `ui/src/lib/characterKitFaceRig.ts`, and +`ui/src/lib/cutoutDialogue.ts`. Persistence: +`app/services/character_kit_library.py`. Cleanup: +`POST /api/v1/character-kits/face-rig/cleanup`. HTTP boundary: +`app/_launch_runtime.py`. + +Related: [3D Video compositor](../3d-video-compositor/HOWUSEIT.md), +[Character Creator orbit](../3d-video-compositor/HOWUSEIT.md#54-hunyuan3d-mesh). + +--- + +## 1. What this system is + +| Tool | Tab | Job | +|---|---|---| +| Character Creator | **Characters** | Identity photo → H3 360° orbit → Hunyuan multi-view **mesh** | +| Character Kit library | **3D Video** | Output-folder-scoped 2D puppet (base pose + mouth overlays + blink) | +| Face Rig | **3D Video** → Paso 2 · Labios / ojos | Generate, clean, place, and review overlays | +| Cutout dialogue | **3D Video** | Held/snap mouth keyframes from known text or speech | +| Recipe runner | **3D Video** | Compiles `dialogueBeats` and only **approved** kit pieces | + +Use a kit when you need a **repeatable graphic character** that can speak with +four mouth sprites. Use Character Creator when you need a **3D mesh**. Use H3 +when you need a **performed** face rather than a paper-cutout flap. + +The compositor does not claim phoneme-perfect lip-sync. Cadence is bounded and +graphic. + +### Output folder versus Workspace collection + +Character Kit routes currently scope a **physical output folder**. Their +`workspace` query/body field retains that name for compatibility; it accepts +`default` or `[A-Za-z0-9][A-Za-z0-9_-]*`. It is **not** a logical Workspace +collection ID. The explicit Workspace collection registry +(`/api/v1/workspace-collections`) groups project, asset, and Production IDs +without creating or moving files. See the +[domain model and asset provenance contract](../development/DOMAIN_MODEL_AND_ASSET_PROVENANCE.md). + +--- + +## 2. Hard limits and persistence rules + +1. **Library file:** `{output-folder}/.character-kit-library-v1.json`. It holds + at most **100** kits, **32** poses per kit, **20 MB** encoded JSON, and **500** + provenance objects per kit. +2. **Compare-and-swap:** every write sends `baseRevision`. Stale clients get + `409` with `code: character_kit_revision_conflict`, plus expected and current + revisions. Reload the library; do not retry the same revision. +3. **Persistent sources only:** `blob:` sources are rejected. Use an + `/api/v1/file/...` or `/api/v1/uploads/...` reference, or a filename in the + output folder. Face Rig handoff also rejects transient browser images. +4. **Review gates:** only `approved` pieces mount or enter Recipe inventory. + Generated Face Rig states, mouth packs, and cleaned overlays remain `pending` + until reviewed and saved. +5. **Mouth states:** `closed`, `small`, `wide`, and `round`. Eye states are + `open` and `blink`; the Face Rig generator calls the first one `open-eyes`. + Unknown keys are rejected with `400`. +6. **Pose-local anchors:** offsets are relative to the character, not the 16:9 + frame. Defaults are mouth `{ offsetX: 0, offsetY: -18, scale: 0.05, + rotation: 0 }` and eyes `{ offsetX: 0, offsetY: -28, scale: 0.12, + rotation: 0 }`. Bounds are offset ±200, scale 0.001–20, rotation ±360. +7. **`lookNotes` is UI-only:** style and trait notes help the current editor + build prompts, but `normalize_character_kit` strips the field on save. +8. **Delete is record-only:** deleting a kit removes its library entry, not its + pose PNGs, cleaned overlays, or scene layers. + +--- + +## 3. Data model + +```text +CharacterKitLibrary { version: 1, revision, activeId, kits{} } + +CharacterKit + id, name, style: cutout | children-illustration | anime-2d + identityReference?, base?, poses{} + mouth { closed?, small?, wide?, round? } + eyes { open?, blink? } + anchors { [poseId]: { mouth, mouthStates?, eyes? } } + provenance[] +``` + +Each asset is `{ id, name, source, kind: image|overlay, alphaStatus, +reviewState, prompt?, model?, workspace? }`. The optional asset `workspace` is +the legacy physical output-folder token, not a Workspace collection ID. + +`alphaStatus` is `unknown`, `transparent`, or `opaque`. An image is considered +transparent when at least 1% of pixels have alpha below 250. + +`mountCharacterKitLayers` parents each approved overlay to its pose, sets +`faceBinding`, and starts with the closed mouth visible. Mounting a kit pose +already present in a scene updates its existing layers instead of adding +duplicates. + +Recipe inventory flattens only approved pieces. The active kit is ordered first +so a large output folder does not evict its complete face set under the global +inventory cap. + +--- + +## 4. Operator workflow + +### 4.1 From Character Creator + +1. Stay on **character** (not object). Capture a turnaround view or upload the + subject. +2. **Create / open CharacterKit Face Rig** stores the handoff in + `sessionStorage` (`hocuspocus:character-kit-face-rig-handoff`) and switches + to **3D Video**. +3. If a kit already uses that source as `base` or `identityReference`, the + editor reopens it. Otherwise it drafts a new kit whose handed-off base pose + is approved for the initial review step. + +Character Creator does not generate visemes. Object mode is rejected with the +message *Face Rig is for Character Kits.* + +### 4.2 From a compositor layer + +1. Open **3D Video** and add a full-body cutout (generated image or transparent + PNG). +2. Choose **New kit from selected base layer**, or assign **Selected → base / + pose / mouth / blink**. +3. Review alpha and approve the pose before Face Rig generation. +4. Open **Paso 2 · Labios / ojos**: choose style and traits, generate missing + states, optionally apply a mouth pack, **Clean**, place, **Lock all mouths**, + then approve each state. +5. **Save kit** (PATCH), then **Mount pose** into the current scene. + +Until **Save kit** is pressed, changes are editor state only; the LLM and other +scenes cannot see pending pieces. + +### 4.3 Face Rig generation + +Prompts are built by `characterKitPosePrompt` and `faceRigPrompt`. The user +supplies style and traits; the helper requests an isolated transparent overlay +or a full-body standing cutout for the pose. + +Image jobs use the selected Studio image model with `strictReference: true` and +the approved pose as identity. The negative prompt forbids a full head/body, +skin rectangle, background, text, glow, and shadow. + +The six generator states are `closed`, `small`, `wide`, `round`, `open-eyes`, +and `blink`. The first four populate `kit.mouth`; the last two populate +`kit.eyes`. Every generated state starts `pending`. + +**Mouth packs** are static assets (no GPU generation): +`GET /character-kit-presets/mouths/manifest.json`. Available packs currently +include `paper-cut`, `children-illustration`, `limited-anime`, `felt-puppet`, +`comic-ink`, and `watercolor`. Applying a pack attaches pending overlays; place +and approve them before mounting. + +**Wipe mouth box** paints an ellipse with nearby samples and uploads a new PNG +named like `--mouthless.png`. It registers a new pose and does +not delete the original. + +**Cleanup** (`POST /api/v1/character-kits/face-rig/cleanup`) runs rembg U2Net +and crop-to-alpha. `padding` is 0–64 (default 8). It writes +`{stem}.cleanup-{8hex}.png` and never overwrites the original. The endpoint +accepts an image inside the uploads root or the selected output folder; a full +body source will be cropped to its opaque bounding box, so cleanup is intended +for one overlay. + +Placement warnings from `assessFaceRigPlacement` never auto-approve. Typical +mouth scale is ≤ 0.12 and blink scale ≤ 0.20. Full-body cutouts usually put the +mouth above the chest (`offsetY` more negative than −8). + +**Lock all mouths** copies one calibrated mouth box to +`closed`, `small`, `wide`, and `round` for that pose. The equivalent eye action +copies the box to `open-eyes` and `blink`. + +### 4.4 Preview speech (Face Rig only) + +`previewFaceRigDialogue` plans **2–4 seconds** of visemes with the same cadence +as scene dialogue. Missing shapes fall back from `wide` to `small`, `round`, or +`closed` as available. The preview does not write scene keyframes. + +--- + +## 5. Cutout dialogue in a scene + +Bind overlays to the selected pose first (`faceBinding` plus a parent +relationship). A speaking kit needs at least one of `small`, `wide`, `round`, +or a legacy overlay identified as `open`; `closed` is optional but provides a +safe resting shape. + +| Action | Result | +|---|---| +| Animate from line | `planCutoutDialogue(text, start, end, fps)` → opacity keyframes | +| Detect from audio | speech units → per-word plans, then the same compiler | + +Cadence in `ui/src/lib/cutoutDialogue.ts`: + +- Minimum hold is `max(2/fps, 0.12s)`. +- Consonants and punctuation map to `closed`; `o/u` to `round`; `a/e` to + `wide`; other vowels to `small`. +- First and last beats settle on `closed` so a cut does not freeze on an open + shape. +- A very short vowel-bearing word gets one centre pulse when cadence allows it. +- Missing speaking sprites fall back to the available open/wide/small/round + layer. + +Persisted scene records use `dialogueBeats[]`: + +```json +{ + "id": "beat-1", + "text": "The square is frozen.", + "start": 0.4, + "end": 2.8, + "mouthLayerIds": ["kit-luma-mouth-wide", "kit-luma-mouth-closed"], + "audioTrackId": "speech-1", + "confidence": "known-text" +} +``` + +`confidence` is `known-text`, `aligned-audio`, or `energy-fallback`. Editing +text, speaker, or timing recompiles mouth keyframes through +`rebuildCutoutDialogueLayers`. Never target the hero/base plate as a mouth +layer. + +Narrative templates with mouth slots include `cutout-talking-head` +(`mouth-open` / `mouth-closed`) and `cutout-speaking-blink`. Hold/snap-only +templates (`cutout-dialogue-hold`, `cutout-reaction-snap`) do not flap a mouth. + +--- + +## 6. Recipes + +The Recipe runner (`ui/src/lib/sceneRecipe.ts`) receives approved kit inventory +tagged `APPROVED_CHARACTER_KIT id=…; role=base|pose/…|mouth/…|eyes/…`. + +Rules given to the planner: + +- Keep body and face pieces from the **same kit ID**. +- Spoken cutout dialogue needs a `speech` audio entry and a top-level + `dialogueBeats` entry with `audioTrackId`, exact text, time range, and + `mouthLayerIds`. +- Multi-shot recipes must set per-shot `audioTrackIds` and `dialogueBeatIds`. + `[]` means silent / no mouths. Do not copy the full mix into every shot. + +`compileRecipeDialogue` turns those beats into ordinary hold keyframes. The +export sidecar stores the compiled scene plus the recipe. + +--- + +## 7. HTTP contract + +These routes are scoped to a physical output folder. The query/body field is +called `workspace` for compatibility. This differs from Director pipeline +routes, which always use the server's active output folder, and from logical +Workspace collection routes. + +Set a base URL for the running HocusPocus instance in the examples below: + +```bash +export HOCUSPOCUS_URL=http://127.0.0.1:7860 +``` + +### `GET /api/v1/character-kits/library?workspace=default` + +Returns the normalized library, or an empty +`{ version: 1, revision: 0, activeId: "", kits: {} }` when the file is missing. + +```bash +curl "$HOCUSPOCUS_URL/api/v1/character-kits/library?workspace=default" +``` + +### `PATCH /api/v1/character-kits/library/kits/{kit_id}` + +Creates or replaces one kit while leaving its neighbours untouched. `baseRevision` +is required by the compare-and-swap contract; `makeActive` defaults to true. + +```bash +curl -X PATCH "$HOCUSPOCUS_URL/api/v1/character-kits/library/kits/luma" \ + -H "Content-Type: application/json" \ + -d '{ + "workspace": "default", + "baseRevision": 0, + "makeActive": true, + "kit": { + "version": 1, + "id": "luma", + "name": "Luma", + "style": "cutout", + "base": { + "id": "luma-base", + "name": "Luma base", + "source": "luma-base.png", + "kind": "image", + "alphaStatus": "transparent", + "reviewState": "approved" + }, + "poses": {}, + "mouth": {}, + "eyes": {}, + "anchors": { + "base": { + "mouth": { "offsetX": 0, "offsetY": -18, "scale": 0.05, "rotation": 0 } + } + }, + "provenance": [] + } + }' +``` + +`400` means the body or kit is invalid. `409` means another client advanced +the library revision; reload before writing again. + +### `DELETE /api/v1/character-kits/library/kits/{kit_id}` + +Uses the same CAS contract. Body: `{ "workspace": "default", "baseRevision": 1 }`. +Returns `404` if the kit ID is missing. Source files are deliberately retained. + +### `POST /api/v1/character-kits/face-rig/cleanup` + +```bash +curl -X POST "$HOCUSPOCUS_URL/api/v1/character-kits/face-rig/cleanup" \ + -H "Content-Type: application/json" \ + -d '{ "workspace": "default", "source": "mouth-wide.png", "padding": 8 }' +``` + +The response includes `filename`, public `source`, `original`, `width`, +`height`, alpha metrics, `method: "rembg-u2net"`, `model: "u2net"`, and +`padding`. `400` / `404` indicate a missing or disallowed image. + +--- + +## 8. Pitfalls + +- Treating Face Rig as a 3D blendshape or H3 speech. It is a small set of PNG + overlays plus opacity holds. +- Mounting before review. The compositor rejects unapproved poses and overlays. +- Saving from two browser tabs with the same revision. The second tab receives + `409` and must reload. +- Mixing kit A's mouth with kit B's body in a hand-edited recipe. Inventory + guidance prevents this, but a manually edited JSON can still be inconsistent. +- Naming overlays without mouth/viseme tokens or `faceBinding`. Discovery has + a legacy label fallback; mounted kits set semantic bindings explicitly. +- Expecting `lookNotes` to survive Save kit. +- Running cleanup on a full-body pose when you intended to clean only one + overlay; the endpoint crops the opaque bounding box. +- Trying Face Rig from Character Creator object mode; it is rejected on purpose. +- Confusing an output-folder token in these routes with a Workspace collection + ID. The latter is metadata and does not select files. + +--- + +## 9. Files to read next + +| Path | Why | +|---|---| +| `ui/src/lib/characterKit.ts` | Types, mounting, and Recipe inventory | +| `ui/src/lib/characterKitFaceRig.ts` | Prompts, packs, anchors, wipe, preview | +| `ui/src/lib/cutoutDialogue.ts` | Viseme planner and keyframe compiler | +| `ui/src/lib/characterKitHandoff.ts` | Creator → 3D Video session handoff | +| `ui/src/lib/sceneRecipe.ts` | `dialogueBeats` and `APPROVED_CHARACTER_KIT` | +| `app/services/character_kit_library.py` | CAS store and validation | +| `app/services/character_kit_face_cleanup.py` | rembg and crop | +| `ui/public/character-kit-presets/mouths/manifest.json` | Pack IDs and files | +| `tests/test_character_kit_library.py` | Server contract | +| `ui/tests/characterKitFaceRig.test.mjs` | Client Face Rig contract | diff --git a/docs/scene-recipe-gap.md b/docs/scene-recipe-gap.md index 541f73d6..06499ad6 100644 --- a/docs/scene-recipe-gap.md +++ b/docs/scene-recipe-gap.md @@ -4,12 +4,21 @@ Field-by-field audit of everything a `Scene` can hold against everything a `SceneRecipe` can express. Six readers swept one domain each, and every claim was then re-checked by a second reader briefed to refute it. +**Status (2026-09-02):** a Scene → Recipe serializer now exists at +`ui/src/lib/sceneToRecipe.ts`. It copies `animation.keyframes`, `strip` / +`seamOccluder`, `relationship`, `visible`, `locked`, `faceBinding`, per-layer +`effects`, and related timing. The counts below are the **T2.3 snapshot** from +before that serializer; do not treat “there is no serializer” as current. +Re-audit before using this file as a blocker list. Character Kit / cutout +dialogue transport is documented in +[`character-kits/HOWUSEIT.md`](character-kits/HOWUSEIT.md). + **122 fields audited. 65 cannot be expressed at all, 28 partially, 29 fully.** 118 claims survived verification; 4 were corrected. -There is no `Scene → Recipe` serializer anywhere in `ui/src`. The sidecar written +At audit time there was no `Scene → Recipe` serializer in `ui/src`. The sidecar written beside every exported MP4 pairs the recipe with the compiled scene and presents -the recipe as the clip's reproduction. For 65 fields that is already untrue, and +the recipe as the clip's reproduction. For 65 fields that was already untrue, and it becomes untrue for a given clip the moment anyone edits it. ## Where the losses cluster diff --git a/docs/workspaces/HOWUSEIT.md b/docs/workspaces/HOWUSEIT.md index e19cec1b..37042768 100644 --- a/docs/workspaces/HOWUSEIT.md +++ b/docs/workspaces/HOWUSEIT.md @@ -1,8 +1,8 @@ # HOWUSEIT — Workspaces tab (Director threads) -The **Workspaces** gallery tab is a **Director generation-thread dashboard**. It is **not** the sidebar switcher that isolates output directories (also called “Workspaces” in the README). +The **Workspaces** gallery tab is a **Director generation-thread dashboard**. It is **not** the sidebar selector for physical output folders, and it is not the explicit Workspace collection registry. -UI tab: **Workspaces** (`mediaFilter: workspaces`). Code: `ui/src/features/workspaces/`. Persistence: `app/services/director_pipeline.py`. HTTP: `app/_launch_runtime.py`. +UI tab: **Workspaces** (`mediaFilter: workspaces`). Code: `ui/src/features/workspaces/`. Persistence: `app/services/director_pipeline.py`. HTTP: `app/_launch_runtime.py`. Domain terminology: [Domain model and asset provenance](../development/DOMAIN_MODEL_AND_ASSET_PROVENANCE.md). Related: [Video Editor / mixes](../video-editor/HOWUSEIT.md), [H3 prompt revisions](../h3-prompt-revisions.md). @@ -12,21 +12,24 @@ Related: [Video Editor / mixes](../video-editor/HOWUSEIT.md), [H3 prompt revisio | Name | What it is | |---|---| -| Output workspace | Isolated save directory (`services.active_workspace`, default `default`). Favorites and outputs are per directory. | -| Workspaces tab | List of saved Director / music-video **pipelines** in the **active** output workspace. Inspect the shot queue, edit prompts, toggle vocal drive, batch-rewrite, resume, rejoin. | +| Output folder | Isolated physical save directory (`services.active_workspace`, default `default`). Favorites and outputs are per directory. The older API often calls this value `workspace`. | +| Workspace collection | Explicit logical collection of project, asset and Production IDs. It is not a directory and does not move files. Its API is `/api/v1/workspace-collections`. | +| Workspaces tab | List of saved Director / music-video **pipelines** in the **active** output folder. Inspect the shot queue, edit prompts, toggle vocal drive, batch-rewrite, resume, rejoin. | -Director pipeline routes **do not** take `?workspace=`. They always read/write the server’s active workspace. +Director pipeline routes **do not** take `?workspace=`. They always read/write the server’s active **output folder**. A logical Workspace collection is metadata only and does not change this routing rule. -Typical flow: generate a song or Director video elsewhere → the thread appears in the left list (newest first) → select it → edit shots → **Start / resume videos** or **Regenerar vídeo completo** (rejoin). +Typical flow: generate a song or Director video elsewhere → the thread appears in the left list (newest first) → select it → edit shots → **Start / resume videos** or **Regenerar vídeo completo** (rejoin). The run may optionally be linked to a Workspace collection, but the files remain in the selected output folder. --- ## 2. Queue inspection -`GET /api/v1/director/pipelines` — saved threads. -`GET /api/v1/director/pipelines/active` — in-memory runs (recovery). +`GET /api/v1/director/pipelines` — saved threads (newest first). +`GET /api/v1/director/pipelines/active` — in-memory runs (recovery). `GET /api/v1/director/pipelines/{pid}` — full state, **hydrated**. +List query: `?limit=&offset=`. **`limit=0` (default) returns every saved pipeline** and parses each JSON. The Workspaces UI pages **8** (`DASHBOARD_PIPELINE_PAGE_SIZE`) and uses `total` plus `loadMorePipelineList` so opening the tab does not hydrate the whole archive. `GET …/{pid}` is still required for the selected thread. Status polls are serialised in the UI; do not fire a full unpaged list on an interval. + Hydration (`hydrate_queue_clips`): | Condition | `queue_source` | @@ -128,7 +131,8 @@ Rejoin uses the mix path in [Video Editor / mixes](../video-editor/HOWUSEIT.md) ## 7. Pitfalls -- Confusing this tab with the **output-directory** switcher. Threads are scoped to whichever directory is active. +- Calling `GET /api/v1/director/pipelines` with the default `limit=0` from a poller. That re-parses every pipeline JSON. Use `limit`/`offset` (the tab uses 8). +- Confusing this tab with the **output-folder** selector or with a logical Workspace collection. Threads are scoped to whichever output folder is active. - Editing prompts while Director is sampling → `409`. - Batch rewrite without a loaded local LLM → `POST /api/v1/llm/generate` fails. - Toggling a shot to mute without cleaning vocal verbs in the visual prose → H3 preflight rejects the shot.