This file tells any coding agent (Claude, Codex, Cursor, Gemini, Copilot, etc.) how to
navigate and contribute to devflow.
devflow is a collection of Agent Skills that
encode a phase-based coding workflow: gather-requirements → create-plan → implement-step →
finalize-feature → review-changes, coordinated by a thin devflow orchestrator skill.
The repo itself is built using the same workflow — it dogfoods its own skills.
-
Do not edit accepted requirements in place. Requirements under
docs/features/<slug>/requirements.mdare immutable once markedstatus: accepted. The only legal way to change prior behavior is to author a new requirement withsupersedes: [REQ-xxxx]and regenerate.devflow/state.yml. -
Run code early and often. Prefer a temp script you delete later over a long speculative edit loop. Temp scripts live under
tmp/and must be cleaned up in thefinalize-featurestep. -
Pause at commit boundaries. Each plan step ends at a commit point. Stop, summarize what was done and why, and wait for the engineer to review before proceeding.
-
If the plan conflicts with reality, stop and escalate. Do not silently work around the plan. Update
plan.md,decisions.md, and any affected entries inscenarios.ymlfirst, then resume. -
SOLID (subset): favor Single Responsibility, Open/Closed, and Dependency Inversion in any code produced by
implement-step. -
Keep each
SKILL.mdunder ~500 lines. Push detail toreferences/. The Agent Skills spec is deliberately designed around progressive disclosure — respect it.
devflow/
├── skills/ # skills installed by the skills CLI
│ ├── devflow/ # orchestrator (Phase 0)
│ ├── gather-requirements/ # Phase 1
│ ├── create-plan/ # Phase 2
│ ├── implement-step/ # Phase 3
│ ├── finalize-feature/ # Phase 4
│ └── review-changes/ # Phase 5
├── examples/
│ └── <sample-feature>/ # dogfood artifacts (requirements / plan / scenarios)
├── AGENTS.md # this file
├── CLAUDE.md # pointer to AGENTS.md for Claude Code
├── README.md # human-facing overview + install
└── LICENSE # MIT
Every skill directory follows the Agent Skills format:
<skill-name>/
├── SKILL.md # required, includes YAML frontmatter (name + description)
├── references/ # optional, loaded on demand
├── scripts/ # optional
└── assets/ # optional
flowchart LR
user["user kicks off work"] --> devFlow["devflow (orchestrator)"]
devFlow --> req["gather-requirements"]
req -->|requirements.md accepted| plan["create-plan"]
plan -->|plan.md + decisions.md + scenarios.yml| impl["implement-step"]
impl -->|pauses at each commit| user
impl --> final["finalize-feature"]
final -->|AGENTS.md updated, tmp cleaned| review["review-changes"]
review -->|readability + security + audits pass| done["done"]
Intent triggers in each skill's frontmatter description route the agent into the right
phase. The orchestrator hands off; it does not re-implement phase behavior.
- Requirements carry monotonic IDs:
REQ-0001,REQ-0002, … - Each requirements file ends with a machine-readable
deltas:YAML block listingadds/modifies/removes/supersedes. .devflow/state.ymlis the accumulated system contract, deterministically rebuilt from the ordered acceptance log. It is checked in so PRs show contract drift.- Before acceptance,
gather-requirementsdry-runs the delta and runs Tier 1 (structural) and Tier 2 (declarative state / budget / dependency) conflict checks. If anything conflicts, the user resolves via amend draft, supersede prior, or reject draft. review-changesruns a state-drift audit that fails hard if the checked-in.devflow/state.ymldisagrees with a re-fold of the acceptance log.
The full rules live in skills/gather-requirements/references/.
Each feature has a scenarios.yml file that lists scenarios as structured entries. Each
entry has a narrative description, traceability tags, and — once Phase 3 runs — a
tests: list pointing to one or more tests (unit, integration, contract, load, smoke,
e2e) in the consumer repo's native test framework. Scenarios are the spec; tests are the
proof. See
skills/create-plan/references/scenarios-schema.md.
Vocabulary v1 (tags live under tags: on each scenario entry):
- Traceability:
req: [REQ-0017, REQ-0018],plan_step: 3,decision: [DEC-0004],owner: <handle> - Lifecycle:
status: spec-only | tests-written | passing | flaky | deferred - Run matrix:
env: [local, ci],browser: [chromium, firefox],platform: [linux]
Scenario-level fields (outside tags:): pause_after, assumes, locked, examples.
This repo is under active development. The planned step sequence:
- Step 1: scaffold
- Step 2:
devfloworchestrator skill - Step 3:
gather-requirementsskill (+ conflict-detection, state-file, supersede-protocol references) - Step 4:
create-planskill (+ scenarios.yml schema references) - Step 5:
implement-stepskill (+ TDD loop, SOLID, pause-points) - Step 6:
finalize-featureskill (+ handoff checklist) - Step 7:
review-changesskill (+ audit machinery, readability, security, review report) - Step 8: CI validation + finalized README catalog + CHANGELOG
- Step 9: Dogfood on a URL-shortener sample under
examples/
Each step ends at a reviewable commit.
Before pushing, run:
npm run check # validate-skills.sh + check-mermaid.mjsCI enforces the same checks — see .github/workflows/ci.yml.
Node version is pinned in .tool-versions (asdf-compatible; nvm use
also respects it).
Last commit on main: 6944f06 — Step 8: add CI validation, finalize README catalog, seed CHANGELOG. Working tree clean.
Next up — Step 9: Dogfood the pipeline on a URL-shortener sample under examples/.
Scope:
- Create
examples/url-shortener/. A toy consumer repo that the skills will operate on. It should be small enough to read in one sitting (a single language, a few handlers, an in-memory or SQLite store) but real enough to exercise every artifact:requirements.md,plan.md,decisions.md,scenarios.yml,.devflow/state.yml,.devflow/log.jsonl. - Walk Phases 1 → 5 on a single feature. The example feature should be modest (e.g. "shorten a URL and look it up") and produce accepted requirements, a plan, scenarios with real tests, passing code, a finalized handoff, and a clean review report. Commit the dogfood artifacts alongside the toy code so readers can see what each phase outputs.
- Exercise the supersede protocol. Add a second requirement that
supersedes:the first (e.g. "return 404 when the short code is unknown, not 500"). This is the single best way to prove the requirements-as-migrations contract works end-to-end; snapshots ofstate.ymlbefore/after should diff cleanly. - Exercise one hard-block audit. Intentionally introduce a state-drift (or
scenarios-coverage, or traceability) violation in a throwaway branch and
record what
review-changesreports. Then fix it. The goal is a worked example of the failure mode, not a permanent bug. - Cross-reference from
README.md. A short "Try it" section linking toexamples/url-shortener/with a one-paragraph tour of what the reader will find there.
Step 9 should NOT modify any skill. If dogfooding surfaces a skill bug, open it as a
new feature via gather-requirements — don't patch skills inside the example's
commit boundary.
Recommended first action for the resuming session: read this file, then
skills/devflow/SKILL.md, then kick off Phase 1 on the URL-shortener feature.
- Start a new feature with the
gather-requirementsskill (dogfood the workflow). - Keep each
SKILL.md≤ 500 lines; push detail toreferences/. - Run
npm run checkbefore committing — it runsskills-ref validateon every skill and parses every```mermaidblock in the repo. CI enforces the same checks.
MIT.