Problem
Currently there is no systematic way to compare different bot configurations (model choice, search depth, calibration alpha, research count). Parameter decisions are made by intuition, not data.
Proposed Solution
Build a backtest harness that:
- Maintains a fixed set of already-resolved Metaculus questions (diverse types: binary, numeric, date, MC)
- Runs the bot with different configs against these questions (without submitting)
- Records per-config metrics: Brier score, log score, relative-to-CP improvement, cost, latency
- Outputs a comparison table and optionally a Pareto frontier (accuracy vs cost)
Key Design Points
- Store resolved questions as JSON fixtures (question text + resolution + CP at access time)
- Support A/B comparison:
python backtest.py --config-a default --config-b deep-research
- Track token usage per run for cost estimation
- Separate "research quality" score (did the research find the key fact?) from "calibration quality" (was the probability well-placed?)
Acceptance Criteria
Problem
Currently there is no systematic way to compare different bot configurations (model choice, search depth, calibration alpha, research count). Parameter decisions are made by intuition, not data.
Proposed Solution
Build a backtest harness that:
Key Design Points
python backtest.py --config-a default --config-b deep-researchAcceptance Criteria