The agent-driven paid-creative workflow that refuses to lie.
Traceable evidence → original or adapted creatives → sealed QA → ads requested PAUSED and live-verified PAUSED — with receipts at every step.
What is this? A production creative pipeline where AI agents do the creative and strategic work (research, hypotheses, copy, scenes, QA judgment) and fail-closed validators enforce the truth (provenance, rights, hashes, timing, PAUSED-only publishing). The human holds the only keys that matter: activation, budget, spend.
Public scope: AGPL-3.0 engine + the fictional sunrise-demo app only.
This repo ships zero real product data, ad accounts, campaign IDs, or private receipts.
Licensed under AGPL-3.0 — run a modified version as a network service? Share your changes.
| 🏃 Install & run the demo | Quick start — clone → venv → preflight → build (sunrise-demo only) |
| 🧭 Understand the design | Why fail-closed? · The 12-stage loop |
| 🧾 See the receipts | What "proof" means here — capability, readiness, PAUSED readback envelopes |
| 🛠 Onboard your own app | docs/GETTING-STARTED.md — copy apps/sunrise-demo.yaml → your slug |
| 📖 Operator handbook | PLAYBOOK.md — the full operating contract, stage by stage |
python3 scripts/forge.py build --app sunrise-demo --batch-id demo-001 --jobs 2That's the actual QA contact sheet the demo produces on a fresh clone: one concept, two markets (en-US, es-MX), transcreated copy (not translated), brand palette from config, safe zones checked, every PNG hashed and sealed. No accounts, no credentials, no network writes. Demo-only — Sunrise Walks is fictional end to end.
Requires Python 3.11+. Image path needs a local Chrome for HTML→PNG. Video path also needs Node 22 + FFmpeg (Remotion).
git clone https://github.com/davidmosiah/creative-forge.git
cd creative-forge
python3 -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
# optional editable install of the CLI: pip install -e .
# optional: export CREATIVE_FORGE_ROOT="$PWD"
# image path
python3 scripts/forge.py preflight --app sunrise-demo
python3 scripts/forge.py build --app sunrise-demo --batch-id demo-001 --jobs 2
# video path (one-time Remotion setup)
cd remotion && npm ci --ignore-scripts && npm run typecheck && cd ..
python3 scripts/forge.py build-video --app sunrise-demo \
--recipe morning-ritual --locale en-US --batch-id demo-001git clone https://github.com/davidmosiah/creative-forge.git
pip install ./creative-forge
# Demo workspace is bundled in the wheel — image preflight/build works without staying in the tree
creative-forge preflight --app sunrise-demo
creative-forge build --app sunrise-demo --batch-id demo-001 --jobs 2The wheel quick start is exercised in CI. PyPI publication is still pending
Trusted Publisher configuration — this README does not claim a live
pip install creative-forge from PyPI until that external state exists.
Remotion video still needs a full checkout (Path A).
On macOS, full Chrome launches are serialized by default because concurrent
ephemeral profiles are unreliable under load. --jobs still controls other
platforms; an operator may explicitly override the cap with
CREATIVE_FORGE_CHROME_MAX_PARALLEL after verifying the local Chrome build.
# 1. Use the contact sheet as an index, open every original PNG, and write
# one note per artifact_key in qa-review.json before sealing approval.
python3 scripts/qa.py approve --report qa/sunrise-demo/demo-001/report.json \
--reviewer you --review-file qa-review.json
# 2. Static dashboard of everything the sealed artifacts prove
python3 scripts/dashboard.py --app sunrise-demo --opensunrise-demo is a fictional product with fictional research so the whole
pipeline runs end to end with zero real data. Every fictional file says so in
its header. Do not commit real product data, ad-account IDs, or private
receipts into this public tree — keep those in your own private workspace and
point CREATIVE_FORGE_ROOT (or copy only sanitized configs) as needed.
See also: CHANGELOG.md · Contributing
Everyone building with agents hits the same wall: the agent says "campaign published!" — and nothing was published. It reports metrics that don't exist. It would happily spend your budget at 3am.
creative-forge inverts the trust model. Agents author; validators verify; humans authorize.
| The agent decides | The validators enforce | Only a human may |
|---|---|---|
| research angles & hypotheses | provenance: every creative cites recorded research | activate an ad |
| copy, scenes, pacing, cultural fit | rights: source, hash, commercial + derivative scope | set or change budget |
| what a metric result means | schema, char limits, safe zones, encoding, timing | spend money |
| QA judgment (it must actually look) | sealed approvals — one changed byte voids them | waive a physical check |
Three rules with no exceptions:
- Truthful lineage, creative latitude. Every concept declares a traceable
lineage_ref. A recipe may separately declare anexecution_reffor a format or audiovisual pattern—even from another lineage. Structural matching applies only when that execution reference iscompetitor_pattern; otherwise the agent is free to invent hook, composition, copy, scenes, pacing, format, and visual language. Validators never score taste. - A local receipt never proves external state. Publishing needs a fresh
capability receipt, a live destination readback, and a byte-canonical
PAUSEDreadback envelope from the platform. AnACTIVEprovider response cannot be masked by local bookkeeping. - Honest metrics. Missing data is
insufficient_data, never zero. ROAS is never invented. Video results carry hook/hold rates so a losing ad tells you where it lost (the first 3 seconds vs. the body).
flowchart LR
A[1 Readiness] --> B[2 Research 360]
B --> C[3 Taxonomy]
C --> D[4 Brief]
D --> E[5 Concepts]
E --> F[6 Production]
F --> G[7 Localization]
G --> H[8 Render]
H --> I[9 Sealed QA]
I --> J[10 Publish PAUSED]
J --> K[11 Test window]
K --> L[12 Learning loop]
L -->|next brief| D
Every stage has a validator; every hand-off is an artifact on disk you can audit later. The full contract lives in PLAYBOOK.md.
The publish path is where agent pipelines usually lie, so it's the most defended surface here. Creating one ad requires, in order:
- Capability receipt (< 60 min old) naming the real discovered platform tools — including a strictly read-only tool, never a create/update/delete or other write operation.
- Readiness receipts from live checks: the store destination is up, the platform has received app events, attribution is mapped. Raw responses are stored with their sha256.
- A manifest bound to the sealed QA matrix — one ad per concept, capped
per ad set,
PAUSEDhardcoded,activation_allowed: false. - A byte-canonical readback envelope per ad, cross-binding provider, tool,
timestamp,
PAUSED, all platform IDs and the artifact hash:
{"binding": {"ad_id": "…", "artifact_sha256": "…", "status": "PAUSED", "…": "…"},
"provider_response": {"id": "…", "status": "PAUSED", "…": "…"},
"schema": "creative-forge/meta-ad-readback@1", "tool": "…", "observed_at": "…"}scripts/publish.py verify-receipt re-validates the whole chain — and even
then, a verified receipt only proves the audit record: a DONE claim
requires the live readback observed in the current run.
| Piece | What it does |
|---|---|
scripts/forge.py |
preflight / build / build-video / prepare-publish |
scripts/qa.py · video_qa.py |
sealed QA: hashes, dimensions, safe zones, per-artifact visual approval |
scripts/research.py · video_mining.py |
research contracts — structure only, never media reuse |
scripts/audiences.py |
fail-closed audience plans (no sensitive-interest targeting, ever) |
scripts/publish.py |
PAUSED-only manifests, capability receipts, canonical readback envelopes |
scripts/experiments.py |
learning loop: metrics provenance, hook/hold rates, next-brief binding |
scripts/host_assets.py |
content-hash static hosting + live byte-for-byte URL verification |
scripts/dashboard.py |
static evidence viewer — reads sealed artifacts only, invents nothing |
remotion/ |
generic video composition (story · portrait · square), mute-safe by contract |
templates/image/* |
HTML creative templates with per-field char limits and declared safe zones |
apps/sunrise-demo.yaml |
the fictional demo app — copy its shape to onboard a real one |
- 321 tests (
python3 -m unittest discover -s tests -v) — validators, receipts, hardening, contracts - Typechecked video (
cd remotion && npm run typecheck) - CI runs the full suite plus a real Remotion render with sealed QA prep
- Fictional demo data only; the repo ships zero real product or ad-account data
- Competitors' public signals only; their art, media, voices and copy are never reused.
- No religious or sensitive-interest ad targeting — context lives in the creative, language and country.
- Agents author freely; validators enforce truth/rights/state; humans authorize activation/budget/spend.
- Every bound on coverage is logged — silent truncation is treated as lying.
PRs welcome — read CONTRIBUTING.md first (rule #1: never weaken a fail-closed gate). Vulnerabilities: use GitHub's private vulnerability reporting — see SECURITY.md.
AGPL-3.0-only (GNU Affero General Public License v3).
- Free to use, study, modify, and redistribute under AGPL-3.0.
- If you run a modified version as a network service (SaaS, hosted API, multi-tenant agent), you must offer the corresponding source to users of that service — the Affero clause is intentional.
- Full text: LICENSE.
