Skip to content

About

Statistical performance regression detector. Mann-Whitney U + Cliff's delta to compare a candidate run against a baseline. CI gate, Markdown/JSON reports, zero dependencies.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

regression-radar

A performance regression detector. Give it a baseline run and a candidate run; it tells you — with statistics, not vibes — whether the candidate is slower, lower-throughput, more error-prone, or more resource-hungry, and by how much.

Pure Python 3.10+. Zero runtime dependencies (the Mann-Whitney U test and Cliff's delta are implemented from scratch). Local-first.

  baseline.json ─┐
                 ├─> regression-radar ─> findings + verdict
  candidate.json ─┘                       (Markdown / JSON)
                                          exit 1 if regressed

Why this exists

"Is the new build slower?" is deceptively hard to answer well. The naive approaches all fail:

  • Comparing averages hides tail regressions — the median can be flat while p99 doubles.
  • Comparing single percentiles is noisy — one slow sample moves p99.
  • Eyeballing two numbers ignores whether the difference is real or just run-to-run variance.

regression-radar uses the statistics that performance engineers actually reach for:

  • Mann-Whitney U test to decide whether two latency distributions genuinely differ. It's non-parametric, so it makes no assumption that latency is normally distributed — and latency is never normally distributed; it's heavy-tailed and right-skewed.
  • Cliff's delta to measure how big the difference is. A p-value tells you whether a difference exists; the effect size tells you whether to care. On large samples everything is "statistically significant" — the effect size is what stops the tool crying wolf.

This is the same rule-engine-with-cited-evidence shape as perf-advisor, and the JSON output is compatible with incident-pilot so a detected regression can be handed off as an incident.

Install

git clone https://github.com/anshikapundeel/regression-radar
cd regression-radar
pip install -e .

Or run without installing:

python3 -m regression_radar compare examples/baseline.json examples/candidate_regressed.json

Requires Python 3.10+. No other dependencies.

Quick start

# A regressed candidate (slower, lower throughput, more errors, memory leak):
regression-radar compare examples/baseline.json examples/candidate_regressed.json

# A clean candidate (statistically indistinguishable):
regression-radar compare examples/baseline.json examples/candidate_clean.json

# Machine-readable output:
regression-radar compare examples/baseline.json examples/candidate_regressed.json --format json

# CI gating: exit 1 if any MEDIUM+ regression is found
regression-radar compare baseline.json candidate.json --fail-on-regression

# List the detectors
regression-radar list-detectors

Sample output:

## 1. Latency regression in http.request_latency_ms

Severity: 🔴 HIGH   Confidence: high   Detector: latency.distribution_shift

Summary. Candidate latency for 'http.request_latency_ms' is significantly
higher (Mann-Whitney p=0, Cliff's d=+0.961 [large]). Median 11.90ms ->
18.90ms (+58.8%); p99 24.00ms -> 48.00ms (+100.0%).

The detectors

Detector Metric kind Catches
latency.distribution_shift distribution Whole-distribution latency regression (Mann-Whitney + Cliff's delta)
latency.percentile_shift distribution Tail-only blowouts where p99 regresses but the median doesn't
throughput.drop counter Throughput regressions (requests/sec, ops/sec)
errors.rate_increase rate Error-rate increases, including zero-to-nonzero new failure modes
resource.growth gauge CPU / memory growth, with a leak-pattern check for monotonic climbs

Each detector lives in its own file under regression_radar/detectors/ and is ~100 lines. Each produces findings with severity, confidence, a concrete suggestion, and cited evidence — the same contract as perf-advisor (findings without evidence are dropped).

The latency detector also surfaces improvements (candidate faster) as INFO findings — seeing the tool report good news builds trust that it isn't only a bearer of bad.

Input format

A run is a label plus named metric series. Each metric has a kind that tells the detectors how to interpret it:

{
  "label": "baseline-v1.2.0",
  "metrics": {
    "http.request_latency_ms": {
      "kind": "distribution",
      "unit": "ms",
      "samples": [11.2, 12.1, 10.8, 13.5, ...]
    },
    "http.throughput_rps": {
      "kind": "counter",
      "unit": "rps",
      "samples": [4820.0]
    },
    "http.error_rate": {
      "kind": "rate",
      "unit": "%",
      "samples": [0.4]
    },
    "proc.rss_mb": {
      "kind": "gauge",
      "unit": "MB",
      "samples": [512, 540, 561, 590, ...]
    }
  }
}

The four kinds:

  • distribution — latency-like; compared with the full distribution test. Needs ≥5 samples (≥20 for the percentile detector).
  • counter — throughput-like; higher is better, a drop is the regression.
  • rate — error-rate-like; lower is better, an increase is the regression.
  • gauge — resource-like (CPU, memory); growth is flagged, with context caveats.

You can pass two files (baseline.json candidate.json) or one file with {"baseline": {...}, "candidate": {...}} via --pair.

See docs/STATISTICS.md for the full explanation of the tests and why they were chosen.

Optional LLM narration

regression-radar ships with no API keys, no default provider, and no outbound calls. The detectors produce the authoritative findings; an LLM, if configured, only rephrases them into a PR-comment-friendly narrative.

export REGRESSION_RADAR_LLM_PROVIDER=ollama
export REGRESSION_RADAR_LLM_BASE_URL=http://localhost:11434
export REGRESSION_RADAR_LLM_MODEL=llama3.1:8b
regression-radar compare baseline.json candidate.json --llm-explain

Also supports openai-compatible providers. Without the env vars the --llm-explain step is skipped and the deterministic report is still produced. The Markdown/JSON report is always the source of truth.

Using it in CI

The intended workflow: run your benchmark/load test on both the baseline and the candidate build, emit two run.json files, then gate the merge:

- name: Performance regression check
  run: |
    regression-radar compare baseline.json candidate.json \
      --fail-on-regression \
      --output regression-report.md
- name: Post report to PR
  if: always()
  run: gh pr comment "$PR" --body-file regression-report.md

--fail-on-regression exits 1 when any finding is MEDIUM or worse, so the job fails and blocks the merge. The report is posted either way.

What this is and isn't

Is:

  • A statistically honest regression detector you can run in CI or locally. The tests it uses are the right ones for performance data, and they're implemented transparently so you can audit them.
  • Composable with the rest of the stack: reads perf-advisor-style JSON, emits incident-pilot-compatible findings.

Isn't:

  • A benchmark runner. You bring the measurements; this analyzes them. Pair it with your existing load tester (wrk, k6, JMH, etc.).
  • A baseline manager. It compares two runs you give it; it doesn't store historical baselines or do drift-over-time tracking. (That's a natural extension — see the roadmap.)
  • A profiler. It tells you that something regressed and by how much; it doesn't tell you where in the code. Use a profiler for that — the findings point you at which metric to profile.
  • Magic for tiny samples. The distribution test declines (returns no finding) below 5 samples per side, because the normal approximation it uses isn't trustworthy there.

Project layout

regression_radar/
  stats.py              Mann-Whitney U, Cliff's delta, percentiles — from scratch
  model.py              Run, MetricSeries, Finding, Comparison
  detectors/            one file per detector
  compare.py            orchestrator: runs detectors, enforces evidence contract
  loader.py             JSON input
  report.py             Markdown + JSON renderers
  llm.py                optional LLM hook; no keys, no defaults
  __main__.py           CLI

tests/                  25 tests: stats properties, each detector, end-to-end
examples/               baseline + regressed + clean fixtures
docs/                   STATISTICS.md, DESIGN.md

Tests

pip install pytest
pytest -v        # 25 tests

CI matrix runs on Python 3.10, 3.11, 3.12.

Roadmap

  • Mann-Whitney U + Cliff's delta from scratch
  • 5 detectors (latency dist, percentile, throughput, error rate, resource)
  • Markdown + JSON reports, --fail-on-regression CI gate
  • Optional LLM narration hook (no keys baked in)
  • 25 unit + end-to-end tests
  • Historical baseline store — track drift across many runs, not just two
  • Permutation test for small samples (where the normal approximation is weak)
  • Direct adapters for wrk/k6/JMH output formats
  • HTML report with distribution plots
  • Hand off detected regressions to incident-pilot automatically

License

MIT.

About

Statistical performance regression detector. Mann-Whitney U + Cliff's delta to compare a candidate run against a baseline. CI gate, Markdown/JSON reports, zero dependencies.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages