50+ agents in daily production · 100+ built over time · running since Nov 2025
I run a production multi-agent AI organization on Claude. It sends email as me, fixes code without me watching, and operates on real money and real hardware, daily. This account is the open half: the pieces I pulled out, cleaned up, and licensed so anyone can run them.
Everything here is bring-your-own-key and self-hostable. More at jddavenport.com.
Product leader turned AI builder. Currently BYU MBA (Strategy & Product Management, 2025–2027) while running the agent system full-time.
Before the MBA: Principal PM at National Grid — grew a fixed-price digital services business from $2M to $20M+ annualized revenue, prototyped an AI-driven dynamic pricing engine (+22% margins), and launched an internal GenAI platform with 3,000+ monthly active users. Earlier: co-founder/CPO at a Web3 fintech startup (raised $1.2M pre-seed, $600K revenue in 90 days), Senior Consultant at Deloitte (Pfizer, GM, Walmart).
Full career and case study: jddavenport.com
agent-safety-case-study is the honest account of running that system: each capability paired with the risk it creates, the control that bounds it, the tradeoff I accepted, and a dated incident where the control was not enough. It includes the review panel whose 3-of-3 quorum turned out to be 1-of-3 with a hardcoded receipt, and what that cost to find.
Three repos extracted from the live production system, positioned around the four things teams building AI products actually need to think about: orchestration, safety, and evaluation.
clawd-agent-os Production framework for orchestrating hierarchical Claude agent fleets. CEO/worker tiers, git-worktree-isolated spawning, a durable bus that raises instead of silently returning empty, single-registry model tiering with a drift guard. 33 tests. Extracted from a 24/7 live system.
ethos-gate Application-layer safety for autonomous agents: PreToolUse/Stop hooks, three-tier credential policy (UNIVERSAL/STRICT/SANDBOX), an irreversibility guard, and a lethal-trifecta auditor that catches the combination-level exposure no per-leg check can see. 62 tests. MIT.
clawd-evals
Behavioral evaluation framework for agentic systems — tests agent conduct, not output quality. Two-gate model promotion bench with no LLM judge, blast-radius-aware eval tiers, prompt injection resistance suite, irreversibility awareness classifier, and over-refusal detection. 57 tests. The LLM-judge problem and why we don't use one is in docs/philosophy.md.
orchestra-agents The orchestration and eval core as installable code: fan-out under bounded concurrency, a durable message bus with a strict status machine and loud dead-lettering, per-dependency circuit breakers, best-of-N with an adversarial judge, and permission tiers from READ to IRREVERSIBLE behind a human-approval gate and a hash-chained audit log. 61 offline tests. The example runs with no API key.
claude-deploy-kit The same discipline packaged for a real deployment: one default-deny policy chokepoint, a destructive-command deny hook that still works under skip-permissions, a fail-closed eval gate, a redacting audit log, and a secrets-provider contract. Every number in the repo is produced by the included tests.
mcp-judge A calibrated LLM-as-judge exposed over MCP, shipped with the eval harness that proves the calibration. It names four judge failure modes, mid-band compression, severity bias, score-to-decision decoupling, and self-inconsistency, and fixes each one rather than asserting the scores are fine.
claude-bug-squash Autonomous bug fixing that out of the box can never merge. Red-to-green reproduction proof, blast-radius caps checked against the actual diff, a never-touch deny list, a three-seat adversarial panel, shadow mode by default, one auto-merge per day, and an explicit arm switch.
recruit-copilot A Claude Code plugin that treats a job search as a verification problem: intake, goals, scouting, tailoring, and a calibrated grading panel with layout and round-trip gates. It does not apply for you. That is the point, not a limitation.
Agents and tooling: fleetwright (self-hosted agent-operations OS) · deep-research-agent (research with an adversarial verification pass) · harnessview (visualize a Claude Code harness and flag broken wiring) · context-kit (personal-context templates and skills) · voiceclaw (voice agent over the phone)
Evals and analysis: vibe-coding-detector · agi-readiness-auditor · resume-grader · ai-spend-tracker · mastery-engine
Self-hosted tools: openbudget · openplaud · byu-outlook-browser-integration
Teaching and product: ai-fluency-trainer · pitchgrade · venture-val · caseprep · acquisitor
Claude (Opus/Sonnet/Haiku fleet) · Python · Next.js · TypeScript · Playwright · Supabase/Postgres · Vercel · GitHub Actions · MCP · LaunchAgents (macOS) · Tailscale
me@jddavenport.com · LinkedIn · jddavenport.com
Licenses vary by repo. Check the LICENSE file in each one. PRs welcome.
