An agent film crew that turns a real near-miss report into a 90-second cinematic safety film β scripted, storyboarded, shot, QC'd, and provenance-signed under a hard render budget.
An agent film studio that turns a company's own near-miss incident report into a 90-second cinematic safety re-enactment β scripted, storyboarded, shot, QC'd, and cut by a Qwen agent crew under a hard render budget managed by a Line Producer agent with a live cost ledger.
A near-miss is the accident's foreshadowing β and foreshadowing is a narrative device. Foreshadow turns one into the other. Drop in Tuesday's incident report, get a film you can show at Thursday's toolbox talk. The demo ledger line β "$2.71 vs $15,000" β is the pitch.
- 8 distinct Qwen Cloud surfaces, one bill. Screenwriter (
qwen3.7-max+ thinking) β Line Producer allocator (qwen3.6-flash) β Art Dept (qwen-image-2.0-pro) β DP (wan2.7-i2v/wan2.6-i2v-flashasync tasks) β QC Critic (qwen3-vl-plus) β Narrator (cosyvoice-v3-plus). - The budget IS the product. The track asks for "maximum output quality under a limited token budget"; the Line Producer makes that constraint the signature feature β a knapsack allocator that tiers every shot by narrative weight, prices every demotion into a regret log, and enforces a hard-cap kill switch at 2.5Γ budget.
- Cryptographic provenance. Every film ships an Ed25519-signed, Merkle-rooted
manifest.
foreshadow verifyre-hashes every artifact and proves invariants I1βI4 β offline, with zero keys.
flowchart LR
CLI["foreshadow CLI<br/>render Β· replay Β· verify Β· bench"] --> P
subgraph P["pipeline worker β Python 3.12, SQLite + fixtures/cache"]
I["ingest β ECIES seal"] --> S["screenplay β shot plan"] --> B["Line Producer<br/>budget knapsack + kill switch"] --> R["render tiers"] --> Q["dailies QC<br/>β€1 retry / demote"] --> M["stitch + publish"]
end
P <--> T{{"Qwen transport"}}
T --> F["FakeQwen<br/>default Β· deterministic Β· no key"]
T --> L["LiveQwen β DASHSCOPE_API_KEY<br/>qwen3.7-max Β· qwen-image-2.0-pro<br/>wan2.7-i2v / wan2.6-i2v-flash Β· qwen3-vl-plus Β· cosyvoice-v3-plus"]
M --> V["Ed25519-signed Merkle manifest<br/>verify: I1βI4, 1-byte tamper fails"]
P --> FC["infra/fc handler β LIVE on Alibaba Function Compute<br/>managed python3.10 Β· /verify Β· /run Β· /health"]
As built = what runs today (offline, keyless, on FakeQwen). The Function Compute handler is now live on Alibaba Cloud (see βοΈ Deployed below); the rest of the deployed topology β Next.js war-room UI, Supabase, OSS render archive β is specified in ../ARCHITECTURE.md / ../SPEC.md and remains pending (see Status).
Foreshadow is deployed live on Alibaba Cloud Function Compute (managed
python3.10 runtime), so a judge can hit the graded pipeline in the cloud with
no local setup:
| Endpoint | What it does |
|---|---|
/health |
liveness β {"status":"ok"} |
/verify |
replays the signed ledger and re-verifies the Ed25519 signature + Merkle chain and invariants I1βI4 in the cloud (byte_identical_cache: true) |
/run?incident=forklift |
replays the full pipeline for a seed incident and returns the ledger ($2.71 spend), budget mix, QC counts, and Merkle root |
curl https://foreshadow-txebjackop.ap-southeast-1.fcapp.run/health
curl https://foreshadow-txebjackop.ap-southeast-1.fcapp.run/verify
curl "https://foreshadow-txebjackop.ap-southeast-1.fcapp.run/run?incident=forklift"The deployed endpoints run the offline-deterministic FakeQwen pipeline β
the same byte-for-byte replayable, signed-ledger path that the tests grade β so
the cloud response is identical to a local foreshadow replay. Live-Qwen
surfaces stay behind DASHSCOPE_API_KEY; a real DashScope call has been smoke-
verified separately (chat surface), but the deployed /run and /verify paths
are the deterministic replay, and no AI-generated video is produced on this path.
Handler + config: infra/fc/ (handler.py, wsgi.py, s.yaml).
python -m venv .venv
./.venv/bin/pip install -e ".[dev]" # includes the offline animatic renderer
./.venv/bin/pytest # -> 421 passed, 100% coverage
./.venv/bin/foreshadow replay --incident forkliftreplay rebuilds the exact demo film from committed fixtures with no network
and no key, then verifies its signed manifest against the committed cache
byte-for-byte. (foreshadow is on the venv path after install; the task's
foreshadow replay --incident forklift works once .venv/bin is active.)
421 tests, all passing at 100% coverage (./.venv/bin/pytest after
pip install -e ".[dev]", fully offline via a session-wide socket guard).
Coverage includes: the Line Producer knapsack math,
budget caps and the 2.5ΓB kill switch; invariants I1βI4 (each failed in
isolation, plus a 1-byte tamper of every committed artifact); schema
validation with reject-retry; Ed25519 sign/verify; ECIES seal/unseal
round-trips; byte-identical replay determinism; the ledger, regret log, and
storage layer; and the full 11-stage pipeline end-to-end for all three seeds.
421 passed
./.venv/bin/pytest # 421 passed, 100% cov, offline
./.venv/bin/python scripts/verify_offline.py # socket-guarded replay + I1βI4, exit 0
./.venv/bin/foreshadow replay --incident forklift # rebuild the demo film, zero keys
./.venv/bin/foreshadow preview --incident forklift # β forklift_animatic.mp4 (real, playable)preview renders a real, playable .mp4 you can open in any player β an
offline storyboard animatic (title cards + the narration script + Ken-Burns
motion, all drawn from the deterministic shot plan). It is not AI-generated
footage: FakeQwen never calls wan, so there is no generated video; the animatic
is clearly stamped as such and exists so a judge has something watchable without a
key. (The wan2.7-i2v path in qwen/live.py builds the real request payload and
runs only with DASHSCOPE_API_KEY.)
Benchmarks (per-surface latency/cost + the $2/$4/$8 budget sweep) live in
docs/BENCH.md; regenerate with foreshadow bench.
Take Qwen Cloud out and Foreshadow is not one integration β it is four vendors, an async render queue, and a cross-vendor cost normalizer, and the single-bill ledger that is the demo becomes impossible.
| # | Qwen surface | Without it |
|---|---|---|
| 1 | qwen3.7-max + thinking (screenplay) |
a separate frontier-LLM vendor + prompt router |
| 2 | structured output JSON schema (ShotPlan / BudgetDecision / QCVerdict) | Zod/retry glue + ~10% parse-failure handling |
| 3 | qwen-image-2.0-pro (character sheet + storyboards) |
a Midjourney/FAL account + style-consistency hacks |
| 4 | wan2.7-i2v / wan2.6-i2v async tasks (hero shots) |
a Runway/Pika vendor + webhook infra |
| 5 | wan2.6-i2v-flash (cheap tier) |
no second price tier β the budget-ladder pitch dies |
| 6 | qwen3-vl-plus (grounded dailies QC) |
a GPT-V-class second vendor just for review |
| 7 | cosyvoice-v3-plus (narration) |
an ElevenLabs subscription |
| 8 | Batch API β50% (storyboard fan-out) | full-price fan-out; the sweep costs 2Γ |
Code citations: agents/screenwriter.py (1, 2), agents/art.py (3),
render/orchestrator.py (4, 5), agents/qc.py (6), render/narrate.py (7),
batch.py (8). All model ids are pinned in config.ALLOWED_MODELS and any
other id raises before a call is built.
Honest limitations: wan clips occasionally drift props between shots (QC
demotes, cannot fix); async tasks have no webhook, so we poll; structured output
needed one reject-retry guard for enum drift. See
docs/friction-log.md.
# ββ Code Quality βββββββββββββββββββββββββββββ
ruff check . # lint
mypy src # type check (advisory)
pytest --cov=src/foreshadow --cov-report=term # unit tests + coverage
# ββ Offline judge-path proof βββββββββββββββββ
python scripts/verify_offline.py # socket-guarded replay + I1-I4
# ββ Security ββββββββββββββββββββββββββββββββββ
pip-audit # dependency vulnerability scan| Layer | Tool | Status |
|---|---|---|
| Code Quality | ruff | β clean |
| Type Checking | mypy | |
| Unit Testing | pytest (421 tests, 100% coverage) | β |
| Offline Judge Proof | scripts/verify_offline.py (socket-guarded) |
β |
| Security (SAST) | CodeQL (python) |
β |
| Security (SCA) | Dependabot (pip + github-actions) + pip-audit |
β |
| Secret Scanning | TruffleHog | β |
| CI/CD | 5-stage GitHub Actions pipeline (Quality β Security β Build β Offline Verify β Deploy Gate) | β |
CI runs .github/workflows/ci.yml on every push/PR to main; CodeQL runs on
its own schedule plus push/PR via .github/workflows/codeql.yml.
./.venv/bin/foreshadow verify \
fixtures/cache/forklift/film.mp4 fixtures/cache/forklift/manifest.jsonRe-hashes every artifact, rebuilds the Merkle root, checks the Ed25519
signature, and evaluates I1βI4 (exit 0 = PASS). The manifest format, Merkle
construction, and threat model are specified in
docs/SPEC-PROVENANCE.md. Signing proves pipeline
integrity, not narrative truth β documented honestly, not hidden.
Demo signer public key (Ed25519): 544aa661bf072330... (derived
deterministically from a public seed so replays are byte-identical; the demo key
proves mechanism, not identity).
Everything above is real and green offline. The following are intentionally scoped for the offline-first build and are not claimed as done:
- Next.js war-room UI β deferred. The CLI + logs are the demo surface. The
UI (agent lanes, live ledger, player) is designed in
UI.mdbut not built. - Live Qwen integration β behind
DASHSCOPE_API_KEY. Chat surfaces are fully implemented on the OpenAI-compatible endpoint and a real DashScope call has been smoke-verified; the image/video/TTS surfaces are payload-complete builders that raiseLiveSurfaceNotVerifieduntil a key is present (src/foreshadow/qwen/live.py). The graded artifact runs onFakeQwen, and no full captured live-Qwen run / AI-generated video is claimed. - Alibaba Function Compute β β
deployed & live. The handler is deployed on
managed
python3.10at https://foreshadow-txebjackop.ap-southeast-1.fcapp.run (/health,/verify,/run) β see βοΈ Deployed above. It serves the same offline-deterministicFakeQwenreplay path; the OSS render archive is still pending (seeinfra/fc/PROOF.md). - Media on the offline path β stubs, plus a watchable animatic. The graded
replay writes deterministic stub media (
film.mp4is a signed edit-list manifest, not decodable video) so replay stays byte-identical on every machine. For a watchable artifact,foreshadow previewrenders a real, playable.mp4storyboard animatic from that same deterministic data (Pillow + imageio-ffmpeg, optional[preview]deps). This is notwan-generated footage β it is honestly an animatic. Real AI video (wan2.7-i2v) runs only behindDASHSCOPE_API_KEY. - x402 pay-per-film API β stretch, not built. Flagged in COMPLEXITY.md as never-claimed-unless-finished.
- PyPI publish β pending. Installable from the repo with
pip install -e.
build/
βββ .github/ CI/CD, CodeQL, Dependabot, community health files
βββ src/foreshadow/ pipeline, agents, crypto, qwen transports, CLI
βββ tests/ 421 offline tests
βββ seeds/ 3 OSHA-300-style incidents (deterministic)
βββ fixtures/cache/ committed replay artifacts (the OSS archive, locally)
βββ scripts/ bench.py Β· verify_offline.py Β· check_submission_readiness.py Β· regen_cache.py
βββ docs/ BENCH.md Β· SPEC-PROVENANCE.md Β· friction-log.md
βββ infra/fc/ handler.py Β· wsgi.py Β· s.yaml Β· PROOF.md (LIVE on Alibaba FC)
βββ .env.example optional DASHSCOPE_API_KEY + runtime path overrides
βββ LICENSE, pyproject.toml
Bug reports, feature ideas, and PRs are welcome β see
.github/CONTRIBUTING.md for the dev setup and the
pre-PR checklist (ruff check ., pytest --cov, scripts/verify_offline.py).
Please also read the
Code of Conduct and the
Security Policy (private disclosure for vulnerabilities).
MIT Β© 2026 Edy Cu.
This project uses Semantic Versioning with fully automated version management driven by Conventional Commits β the version is never edited by hand.
| Commit type | Bump | Example |
|---|---|---|
fix: β¦ |
patch | 1.0.0 β 1.0.1 |
feat: β¦ |
minor | 1.0.0 β 1.1.0 |
feat!: β¦ or BREAKING CHANGE: footer |
major | 1.0.0 β 2.0.0 |
python-semantic-release keeps the version in sync
across pyproject.toml and src/foreshadow/__init__.py.
- In CI/CD: Stage 6 of the pipeline (
.github/workflows/ci.yml) runs on every push tomain, computes the next version from the commits since the last tag, then commits + tags it automatically. - Locally:
pip install -e ".[release]" semantic-release version # compute + apply the next version and tag