Start with evaluate-a-change.
It is the smallest complete path: cases in, scores out.
Install and build from the repository root, then run the first offline example:
pnpm install
pnpm build
pnpm exec tsx examples/evaluate-a-change/index.ts| Goal | Example | Requirements |
|---|---|---|
| Score one change on the same cases | evaluate-a-change |
Offline |
| See the case grid before you pay for it | plan-before-you-spend |
Offline |
| Wrap an existing agent | foreign-agent-quickstart |
Offline |
| Evaluate several attempts per case | multi-shot-optimization |
Offline |
| Apply a release rule without any search | held-out-gate |
Offline |
| Load cases from folders on disk | eval-fixtures-quickstart |
Offline |
| Record and compare scores over time | scorecard |
Offline |
| Run the same cases across several profiles | profile-matrix |
Offline |
| Goal | Example | Requirements |
|---|---|---|
| Improve with your own candidate generator | selfimprove-quickstart |
Offline |
| Declare source units, audit a checker, and retain fresh final evidence | evaluation-integrity |
Offline; build the package first |
| Improve one prompt with official GEPA in one call | self-improve-optimizer |
Python GEPA package and an LLM endpoint |
| Let a metered coding agent drive the optimization | agent-engine-optimizer |
Python GEPA package, the claude CLI, and an LLM endpoint |
| Let another package own the text search | adapt-a-text-optimizer |
Offline |
| Compare official GEPA and SkillOpt | compare-optimization-methods |
Python optimizer packages and an LLM endpoint |
Use the comparison guide for installation and running selected methods. It also documents endpoint settings, rates, execution owners, and GEPA recipes.
| Goal | Example | Requirements |
|---|---|---|
| Register the rules before the data arrives | sealed-experiment |
Offline |
| Certify a result that has no answer key | verify-without-an-answer-key |
Offline |
| Goal | Example | Requirements |
|---|---|---|
| Get a report from runs you already have | analyze-existing-runs |
Offline |
| Get cited findings out of a failed batch | custom-trace-analyst |
Offline |
| Analyze human approvals and rejections | customer-feedback-loop |
Offline |
| Analyze OpenTelemetry spans | customer-otel-traces |
Offline |
| Goal | Example |
|---|---|
| Run public benchmark adapters | benchmarks |
| Compare optimizers on AppWorld tasks | AppWorld |
| Export supervised and preference rows | publish-rl-dataset |
| Fine-tune through Prime Intellect | fine-tune-with-prime-rl |
The AppWorld comparison requires separate AppWorld and optimizer Python environments, plus an LLM endpoint.
The GSM8K comparison reads a local dataset file from AGENT_EVAL_GSM8K_PATH.
Produce it from the GSM8K test split with Python and datasets:
mkdir -p ~/.cache/agent-eval
python -c "from datasets import load_dataset; import json; \
[print(json.dumps({'id': f'gsm8k-test-{i}', 'question': r['question'], 'answer': r['answer']})) \
for i, r in enumerate(load_dataset('openai/gsm8k', 'main', split='test'))]" \
> ~/.cache/agent-eval/gsm8k.jsonlor, without Python, from the upstream source of the Hugging Face dataset:
mkdir -p ~/.cache/agent-eval
curl -L https://raw.githubusercontent.com/openai/grade-school-math/master/grade_school_math/data/test.jsonl \
| jq -c '{id: ("gsm8k-test-" + (input_line_number | tostring)), question, answer}' \
> ~/.cache/agent-eval/gsm8k.jsonlEach output line holds {id, question, answer} — the exact shape benchmarks/gsm8k/index.ts loads.
| Goal | Example |
|---|---|
| Coordinate workers across processes | distributed-driver |
| Run setup, execution, and scoring in one work directory | same-sandbox-harness |
| Receive shipped search ledgers and trace spans | hosted-ingest-server |
_shared/ holds fixtures reused by several examples.
It is not a standalone example.