Skip to content

Build backtest harness for parameter tuning #108

Description

@RuizhangZhou

Problem

Currently there is no systematic way to compare different bot configurations (model choice, search depth, calibration alpha, research count). Parameter decisions are made by intuition, not data.

Proposed Solution

Build a backtest harness that:

  1. Maintains a fixed set of already-resolved Metaculus questions (diverse types: binary, numeric, date, MC)
  2. Runs the bot with different configs against these questions (without submitting)
  3. Records per-config metrics: Brier score, log score, relative-to-CP improvement, cost, latency
  4. Outputs a comparison table and optionally a Pareto frontier (accuracy vs cost)

Key Design Points

  • Store resolved questions as JSON fixtures (question text + resolution + CP at access time)
  • Support A/B comparison: python backtest.py --config-a default --config-b deep-research
  • Track token usage per run for cost estimation
  • Separate "research quality" score (did the research find the key fact?) from "calibration quality" (was the probability well-placed?)

Acceptance Criteria

  • 50+ resolved questions across binary/numeric/date/MC types
  • CLI that runs bot against fixtures and outputs score table
  • At least one A/B comparison documented in repo

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions