- Response language follows
languagesetting in.agents/oma-config.yamlif configured. - Follow
.agents/skills/_shared/core/execution-policy.mdfor authorization, clarification, verification, and completion. Execute required steps on the selected path in dependency order; apply documented branch and skip conditions. - Key-optional by default (backend rule 11). The baseline path uses zero external API keys: the agent writes the script, oma-voice does TTS, oma-image does visuals, captions are key-free, Remotion composites. Every paid upgrade (Pexels stock, Pixelle AIGC) sits behind a key-free fallback and is off by default. Never disable the fallback to force a real call, and never silently drop a requested real path.
- Determinism boundary =
render-spec.json+ asset files (+ seed + embedded Pretendard). "Reproducible from script/assets, not from brief." Never edit assets or render-spec by hand after generation; re-run the stage that produced them. - Demo capture is human-supplied.
--source web --urlgives URL context only; it neither opens nor records a browser. The human records the flow with Cap or another recorder, then supplies--capture <path>. Never automate login or capture credentials. Mask URL query tokens in logs and manifests. - The
oma videoCLI owns the pipeline. This workflow owns the brief, the agent-authored script, the QA loop, and decision checkpoints. Do NOT reimplement orchestration, provider selection, or rendering in the workflow. - Code intelligence and state are separate capabilities. Follow
.agents/skills/_shared/core/code-intelligence.mdfor configured tools and native fallback. Follow.agents/skills/_shared/runtime/memory-protocol.mdfor run tracking and result paths. Do not require an MCP memory tool or automatically install, initialize, or track a repository. - Read the oma-video skill BEFORE starting. Read
.agents/skills/oma-video/SKILL.mdand follow its Core Rules and execution protocol, includingresources/execution-protocol.md. If the skill is not installed, stop and ask the user to runoma installfirst.
Vendor note: This workflow executes inline. The script-authoring step (Step 3) is performed by the running agent itself — the agent is the LLM key (agent-as-key). A subagent may be spawned only for broad multi-scene research; the asset pipeline (Steps 4-7) runs through the
oma videoCLI, not through subagents.
L1 Decision Events
Emit required L1 decisions by calling oma state emit directly, as documented in .agents/skills/_shared/runtime/event-spec.md.
This workflow has two required checkpoints: mode-selection (Step 2) and cost-confirmation (Step 5). Do not skip either emit/verify pair.
Mode Routing
Resolve the mode first — it determines aspect, source, visual track, and compositor. If the user did not name a mode, infer from intent and resolve material ambiguity at Step 2.
| mode | aspect | source | visual track (default → opt) | compositor | output |
|------|:---:|------|------|------|------|
| shorts | 9:16 | synthetic (topic → clip) | oma-image stills · Pexels (opt) · Pixelle AIGC (opt) | Remotion · MPT alt | shorts-<slug>.mp4 |
| explainer | 16:9 / 9:16 | README · code · data | oma-slide frames + oma-image diagrams + code | Remotion (deterministic) | explainer-<slug>.mp4 |
| demo | 16:9 | Human recording via --capture; --source web --url adds URL context | raw footage (default) · Remotion intro · zoom · callouts (--polish) | Remotion polish | demo-<slug>.mp4 |
Intent heuristics: "reel / TikTok / short / hook" → shorts; "walkthrough / how it works / from the README / explain the architecture" → explainer; "record / screen / show the app running / product demo" → demo.
For demo, resolve the recording path first. Use --source file --capture <path> for a recording; --source web --url <url> --capture <path> adds URL context to that human recording. The CLI does not open or record a browser. Raw footage is the default output; --polish overlays the Remotion Demo composition.
Cost Guardrail & Key-Optional Notes (read before Step 4)
-
Guardrail: default
cost.guardrail_usd: 0.20in.agents/skills/oma-video/config/video-config.yaml(reused from oma-image). Any provider whose estimated cost meets or exceeds the guardrail requires spend authorization (-y/--yesor the Step 5 checkpoint). Reuse an existing authorization covering that provider and amount.--max-usd <n>overrides the threshold. -
Key-optional pairs (real path is gated; fallback is always wired):
| capability | real (key/resource) | key-free fallback | deferred marker | |------|------|------|------| | stock video | Pexels (
PEXELS_API_KEY) | oma-image stills + Ken Burns |TODO(oma-deferred): pexels| | AIGC video | Pixelle-MCP + RunningHub (RUNNINGHUB_API_KEY) | oma-image stills |TODO(oma-deferred): pixelle| | caption timing | voicebox-stt (MCPvoicebox_transcribe→ REST) | estimate |TODO(oma-deferred): whisper-cpp| | music mixing | Strudel offline render (oma video doctor --install-strudel) | render without music | — | | premium TTS | (not needed — oma-voice is local) | — | — | -
Pixelle AIGC is a community MCP: off by default, requires one-time explicit user consent plus a source review before connecting, and is always cost-gated on RunningHub credits.
-
Asset-provider fallbacks are not failures. A run may use key-free timing, visual, caption, or music paths and record
pathTaken: fallbackwith a warning. An unavailable, failed, or invalid compositor remains an error.
Step 1: Resolve Brief & Preflight
- Capture the brief from the user's request. If absent, ask:
What is the video about? Give me a one-line brief, and a mode if you have one (shorts / explainer / demo). - Run the readiness check and surface gaps before spending any time on assets:
This reports Node / Chromium / FFmpeg, the Remotion project, the embedded Pretendard font, Voicebox MCP (oma-voice), oma-image vendors, optional Pixelle-MCP, Cap, and MPT readiness. Doctor does NOT auto-bootstrap — plainoma video doctor --output jsonoma video doctoronly reports. If Remotion is not yet installed, runoma video doctor --install(one-time: deps + Chrome Headless Shell + Pretendard font fetch) — do not install during a run. MPT needs a one-timeoma video doctor --install-mpt(clone + venv + deps). - If doctor reports a hard blocker for the chosen mode (e.g. no compositor for
shorts/explainer), report the remediation and stop. If only an optional provider is missing (Pexels, Pixelle, Cap), note it and continue on the fallback. - Record run start in the configured file-memory path: brief summary, requested mode, and doctor result.
Step 2: Confirm Mode & Plan
- State the resolved mode, aspect, locale, caption style, visual track, and compositor you intend to use, and the expected output name.
- For
demomode, state up front: "Capture is performed by a human." Resolve the source:--source file: require--capture <path>; if absent, ask the user to record and provide the file path before proceeding.--source web --url <url>: also require--capture <path>. The URL provides context only; the human records the flow separately. No login or browser capture is automated, and URL query tokens are masked in logs and the manifest.
- Apply
.agents/skills/_shared/core/execution-policy.md: proceed when the requested work or decision is already authorized; ask only for a material missing decision or new authorization. - Once the mode is resolved under the execution policy, emit and verify the mode-selection decision with its actual authorization source:
oma state emit "decision.made" '{"subject":"video.mode-selection","decision":"<resolved mode and pipeline plan>","rationale":"<existing instruction, delegated choice, or new user decision authorizing the plan>"}' oma state verify --workflow video --checkpoint mode-selection
Step 3: Author script.json (agent-as-key)
The agent writes the script — this is the start of the determinism boundary. Do NOT call an external LLM; you are the script provider.
- Produce a script honoring the
script.jsonschema (schemaVersion: "1.0"— required literal — plusmode, aspect, locale, title, scenes[{id, durationSec, narration, onScreenText, visual{kind: still|clip|mixed|slide|capture, prompt, ref, source}, transition}], music, brand). Author against.agents/skills/oma-video/resources/script-schema.md(full field reference + example) — a schema mismatch is exit 4. - Respect limits from
.agents/skills/oma-video/config/video-config.yaml(max_duration_sec: 180,max_scenes: 40). Keep narration tight and per-scene so scene boundaries map cleanly to TTS timing. - Mode-specific sourcing:
shorts: a hook-first synthetic script from the topic; each scene gets avisual.promptfor oma-image.explainer: ground scenes in the README / code / data the user pointed to; mark scenes that should become oma-slide frames vs oma-image diagrams.demo: narration + on-screen callouts over the captured footage; visual refs point at the ingested capture segments.
- Translate narration / on-screen text via oma-translation when
localediffers from the source language (key-free). If oma-translation is absent, keep the source text and let the run warn. - Write the agent-authored script to a file and hand it to the CLI via
--script <path>so it validates against the schema.--scriptis mandatory for the agent-as-key path: without it the CLI builds its own skeleton script from the brief and your authored script is never used. Use--dry-runfor the first pass so the pipeline emitsscript.json+render-spec.json+manifest.jsonwithout rendering:oma video generate "<brief>" --mode <mode> --aspect <aspect> --locale <lang> \ --captions <tiktok|lower-third|none> --visual <auto|generate|stock|aigc|slide> \ --voice <profile|none> --music <upbeat|calm|cinematic|lofi|piano|none> --duration <sec|auto> \ --compositor <remotion|mpt> --seed <n> \ --script <path-to-agent-authored-script.json> --dry-run --output json - Review the emitted
script.jsonfor scene count, durations, and narration quality. Iterate here — fixing the script is cheap; fixing a render is not.
Step 4: Parallel Asset Generation (voice / visual / caption)
The CLI orchestrator fans out the asset tracks per the asset bus. Trigger the full (non-dry) run; the orchestrator runs the tracks and writes them into the run directory. Do not author assets by hand.
oma video generate "<brief>" --mode <mode> [same flags as Step 3, incl. --script <path>, without --dry-run] --output json
The three tracks (per .agents/skills/oma-video/SKILL.md and its execution protocol):
- Voice (oma-voice / Voicebox MCP) → a single
audio/narration-01.wav(all scene lines joined into one track — not per-scene files) +timing.json. Timing source:voicebox-stt(MCPvoicebox_transcribe, REST/transcribefallback, on the generated wav) →estimated(thetts-native/whisper-cppsource values are reserved but deferred). The default voice isnone→ a silent video with estimated timing; pass--voice <profile>for narration. If oma-voice is down, the run falls back to silent + estimated timing and warns — it does not hard-fail. - Visual (per-scene, fallback chain
oma-image → pexels → pixelle) →visuals/scene-NN.*. Default is key-free oma-image stills (aspect snapped to the nearest 16-multiple; Remotion crops to exact frame).--visual stockengages Pexels only whenPEXELS_API_KEYis set;--visual aigcengages Pixelle only after consent + cost gate. Each scene that falls back is recorded withpathTaken: fallback. - Caption (key-free) →
captions.srt/.vtt, aligned totiming.json, styledtiktokorlower-third, with platform safe-area presets. Non-source locales translate via oma-translation; if absent, captions keep the source locale and warn.
Report which path each track took (real vs fallback) and surface any warnings.
Demo capture track (--mode demo)
For demo, the orchestrator produces the footage in place of synthetic visuals, dispatched on --source:
--source file --capture <path>— ingest the human recording (absolutized,$PWD-guarded, format-validated). Without--capture, return the guided protocol and stop.--source web --url <url> --capture <path>— ingest the same human recording and use the URL as context. The CLI does not open a browser, wait on selectors, record the screen, or accept interactive stop controls. The URL and query tokens are masked in logs andmanifest.json.
Step 5: Cost Gate & render-spec.json
- Inspect the cost estimate the orchestrator computed across providers (
cost.usd+ breakdown in the manifest/JSON output). - If the estimate meets or exceeds the guardrail (default $0.20, or
--max-usd), present the breakdown and reuse existing spend authorization if it covers the provider and amount. Otherwise obtain authorization before the paid render proceeds. Then emit and verify the actual decision:
If the user declines, re-run with the key-free providers (dropoma state emit "decision.made" '{"subject":"video.cost-confirmation","decision":"Proceed with the estimated paid cost or fall back to the key-free path.","rationale":"Estimated cost crossed the guardrail; the user confirmed spend or chose the fallback."}' oma state verify --workflow video --checkpoint cost-confirmation--visual stock|aigc) — the fallback chain keeps the run alive. - If the estimate is under the guardrail, note "cost under guardrail ($X.XX < $0.20)" and continue without a confirmation prompt.
- Confirm
render-spec.jsonwas written. This is the deterministic compute boundary:compositor, composition, fps, dimensions, durationInFrames, audio, scenes[], captions, background, seed. The seed is embedded so re-renders are byte-identical.
Step 6: Composite (Remotion — you author the composition; MPT is an alternative)
- Remotion (default, all modes) — oma ships no composition code; you write it per run on the always-latest Remotion:
oma video generatealready scaffolded<runDir>/remotion/(warningcomposition pending). If not, or to refresh:oma video compose <runDir> --output json.- Read, in order:
<runDir>/remotion/AUTHORING.md(contract for this spec), theremotion-best-practicesandremotion-markupSKILL.md paths it lists (remotion-dev/skills at HEAD;remotion-captionswhencaptions.style !== "none",remotion-multimediafor video/audio), and.agents/skills/oma-video/resources/remotion-authoring/<mode>.md. - Write
<runDir>/remotion/src/Root.tsx(+src/components/*): one<Composition id={composition}>consumingrender-spec.json,calculateMetadatafrom props, deterministic (no network/randomness), Pretendard viastaticFile("fonts/PretendardVariable.woff2"). Never edit the generated files. oma video render <runDir> --output json— typecheck →npx remotion render→ ffprobe video-stream/duration validation. A non-zero exit or invalid output is a render failure; use diagnostics to classify toolchain, runtime, or composition causes, then fix and re-render. Use at most three render attempts within ten minutes; then report the diagnostics and recovery artifacts.
- Demo raw vs
--polish: fordemo, the default is the raw captured footage copied through as the output.--polishmeans you author theDemocomposition (intro / callouts / zoom over the capture asbackground).
- MoneyPrinterTurbo (
--compositor mpt, shorts e2e alternative): the agent-written script is injected in custom-script mode; provider keys are env-only and masked in logs. It needsoma video doctor --install-mptonce. Setup or render failures fail with diagnostics; onlyOMA_VIDEO_MOCK=1tests may create a placeholder file. - If the toolchain cannot be fetched (offline, nothing cached):
oma video doctor --installonce online. Do not pin or hand-install Remotion. - Confirm the output MP4 exists, matches
<mode>-<slug>.mp4, has a video stream, and has a positive ffprobe duration.
Step 7: QA Loop
Review the finished video against the brief and the quality bars. Iterate by re-running the smallest upstream stage that owns the defect — never patch the artifact.
- Checklist (priority order: correctness → sync → readability → polish):
- Output plays; duration matches the script total within tolerance.
- Narration audio is present (or intentionally silent) and aligns to scenes.
- Captions are synced to
timing.json, within the safe area, and legible (static windowed cues, CSS-wrapped, Pretendard, design rule 2). - Visuals match each scene's intent; any placeholder visual asset is declared in warnings, and a real validated MP4 remains required.
- Aspect / dimensions are correct for the mode; branding applied as requested. (A requested music mode yields
music/bgm.wavmixed at −18 dB, or a fallback warning and a silent render when Strudel is not installed.)
- Route each defect to its stage:
- script/narration/scene-count → Step 3 (re-author script).
- audio/timing → Step 4 voice track (check oma-voice, re-synthesize).
- wrong/placeholder visual → Step 4 visual track (adjust prompt or
--visualmode). - missing/incomplete demo capture → Step 4 demo capture track (obtain a new human recording and pass
--capture;--urlis context only). - caption sync/wrap/locale → Step 4 caption track (or oma-translation).
- layout/transition/crop → Step 6 edit the composition (
<runDir>/remotion/src) or the render-spec →oma video render(fordemo, toggle--polish).
- Determinism guard: when validating reproducibility, run the golden harness — render-spec and assets must be byte-identical:
OMA_VIDEO_MOCK=1 oma video generate "<brief>" --mode <mode> --seed <n> --dry-run --output json - Stop after three render attempts or ten minutes of render recovery. Report diagnostics and recovery artifacts; if a further change needs a different visual track, mode framing, or compositor, obtain the user's direction before continuing.
- Repeat until the checklist passes or the user accepts the result.
Step 8: Output & Manifest
- Confirm the run directory is complete (mirrors
.agents/results/videos/<runId>-<mode>/):script.json · timing.json · render-spec.json audio/narration-01.wav # single narration track (all lines joined) visuals/scene-*.{jpg,png,mp4} # synthetic modes capture.mp4 # demo: human recording ingested by the CLI captions.srt (+ .vtt) <mode>-<slug>.mp4 manifest.json - Verify
manifest.jsonis the reproducibility record:runId, mode, providers{...}, assets[{path,sha256,bytes,seed}], outputs{video,durationSec,sha256}, cost{usd,breakdown}, warnings[], exitCode. All external assets are copied into the run dir and hashed — no URL refs. Fordemo --source web, the manifest records a masked URL context (query tokens stripped), never credentials. - Report to the user:
- Output MP4 path (absolute) and duration.
- Providers used per track, and which tracks took the fallback path.
- Final cost (
$0.00on the all-key-free path). - Any warnings (silent audio, source-locale captions, placeholder visuals).
- Reproduce command:
oma video render <runDir>.
- Record run completion in the configured file-memory path: run dir, output path, providers, cost, and warnings.
Exit Codes (aligned with oma search fetch)
0 ok · 1 generic · 2 safety · 3 not-found · 4 invalid-input · 5 auth-required · 6 timeout.
Common error → action map:
| error | exit | action |
|------|:---:|------|
| ProviderUnavailableError | 5 | a required provider is down → run oma video doctor, fix or fall back |
| CompositorBootstrapError | 1 | Remotion not installed → oma video doctor install-once, then re-render |
| CostGuardrailError | confirm | estimate crossed guardrail → Step 5 confirmation or drop paid providers |
| CaptureRequiredError | guided | demo needs a human recording → provide --capture <path>; --source web also requires --url context |
| SchemaValidationError | 4 | script/render-spec invalid, missing demo capture, or --source web without --url → fix in Step 3 and re-validate with --dry-run |