A performance regression detector. Give it a baseline run and a candidate run; it tells you — with statistics, not vibes — whether the candidate is slower, lower-throughput, more error-prone, or more resource-hungry, and by how much.
Pure Python 3.10+. Zero runtime dependencies (the Mann-Whitney U test and Cliff's delta are implemented from scratch). Local-first.
baseline.json ─┐
├─> regression-radar ─> findings + verdict
candidate.json ─┘ (Markdown / JSON)
exit 1 if regressed
"Is the new build slower?" is deceptively hard to answer well. The naive approaches all fail:
- Comparing averages hides tail regressions — the median can be flat while p99 doubles.
- Comparing single percentiles is noisy — one slow sample moves p99.
- Eyeballing two numbers ignores whether the difference is real or just run-to-run variance.
regression-radar uses the statistics that performance engineers actually reach for:
- Mann-Whitney U test to decide whether two latency distributions genuinely differ. It's non-parametric, so it makes no assumption that latency is normally distributed — and latency is never normally distributed; it's heavy-tailed and right-skewed.
- Cliff's delta to measure how big the difference is. A p-value tells you whether a difference exists; the effect size tells you whether to care. On large samples everything is "statistically significant" — the effect size is what stops the tool crying wolf.
This is the same rule-engine-with-cited-evidence shape as
perf-advisor, and
the JSON output is compatible with
incident-pilot
so a detected regression can be handed off as an incident.
git clone https://github.com/anshikapundeel/regression-radar
cd regression-radar
pip install -e .Or run without installing:
python3 -m regression_radar compare examples/baseline.json examples/candidate_regressed.jsonRequires Python 3.10+. No other dependencies.
# A regressed candidate (slower, lower throughput, more errors, memory leak):
regression-radar compare examples/baseline.json examples/candidate_regressed.json
# A clean candidate (statistically indistinguishable):
regression-radar compare examples/baseline.json examples/candidate_clean.json
# Machine-readable output:
regression-radar compare examples/baseline.json examples/candidate_regressed.json --format json
# CI gating: exit 1 if any MEDIUM+ regression is found
regression-radar compare baseline.json candidate.json --fail-on-regression
# List the detectors
regression-radar list-detectorsSample output:
## 1. Latency regression in http.request_latency_ms
Severity: 🔴 HIGH Confidence: high Detector: latency.distribution_shift
Summary. Candidate latency for 'http.request_latency_ms' is significantly
higher (Mann-Whitney p=0, Cliff's d=+0.961 [large]). Median 11.90ms ->
18.90ms (+58.8%); p99 24.00ms -> 48.00ms (+100.0%).
| Detector | Metric kind | Catches |
|---|---|---|
latency.distribution_shift |
distribution |
Whole-distribution latency regression (Mann-Whitney + Cliff's delta) |
latency.percentile_shift |
distribution |
Tail-only blowouts where p99 regresses but the median doesn't |
throughput.drop |
counter |
Throughput regressions (requests/sec, ops/sec) |
errors.rate_increase |
rate |
Error-rate increases, including zero-to-nonzero new failure modes |
resource.growth |
gauge |
CPU / memory growth, with a leak-pattern check for monotonic climbs |
Each detector lives in its own file under regression_radar/detectors/
and is ~100 lines. Each produces findings with severity, confidence,
a concrete suggestion, and cited evidence — the same contract as
perf-advisor (findings without evidence are dropped).
The latency detector also surfaces improvements (candidate faster) as INFO findings — seeing the tool report good news builds trust that it isn't only a bearer of bad.
A run is a label plus named metric series. Each metric has a kind
that tells the detectors how to interpret it:
{
"label": "baseline-v1.2.0",
"metrics": {
"http.request_latency_ms": {
"kind": "distribution",
"unit": "ms",
"samples": [11.2, 12.1, 10.8, 13.5, ...]
},
"http.throughput_rps": {
"kind": "counter",
"unit": "rps",
"samples": [4820.0]
},
"http.error_rate": {
"kind": "rate",
"unit": "%",
"samples": [0.4]
},
"proc.rss_mb": {
"kind": "gauge",
"unit": "MB",
"samples": [512, 540, 561, 590, ...]
}
}
}The four kinds:
distribution— latency-like; compared with the full distribution test. Needs ≥5 samples (≥20 for the percentile detector).counter— throughput-like; higher is better, a drop is the regression.rate— error-rate-like; lower is better, an increase is the regression.gauge— resource-like (CPU, memory); growth is flagged, with context caveats.
You can pass two files (baseline.json candidate.json) or one file
with {"baseline": {...}, "candidate": {...}} via --pair.
See docs/STATISTICS.md for the full explanation
of the tests and why they were chosen.
regression-radar ships with no API keys, no default provider, and no outbound calls. The detectors produce the authoritative findings; an LLM, if configured, only rephrases them into a PR-comment-friendly narrative.
export REGRESSION_RADAR_LLM_PROVIDER=ollama
export REGRESSION_RADAR_LLM_BASE_URL=http://localhost:11434
export REGRESSION_RADAR_LLM_MODEL=llama3.1:8b
regression-radar compare baseline.json candidate.json --llm-explainAlso supports openai-compatible providers. Without the env vars the
--llm-explain step is skipped and the deterministic report is still
produced. The Markdown/JSON report is always the source of truth.
The intended workflow: run your benchmark/load test on both the
baseline and the candidate build, emit two run.json files, then gate
the merge:
- name: Performance regression check
run: |
regression-radar compare baseline.json candidate.json \
--fail-on-regression \
--output regression-report.md
- name: Post report to PR
if: always()
run: gh pr comment "$PR" --body-file regression-report.md--fail-on-regression exits 1 when any finding is MEDIUM or worse, so
the job fails and blocks the merge. The report is posted either way.
Is:
- A statistically honest regression detector you can run in CI or locally. The tests it uses are the right ones for performance data, and they're implemented transparently so you can audit them.
- Composable with the rest of the stack: reads perf-advisor-style JSON, emits incident-pilot-compatible findings.
Isn't:
- A benchmark runner. You bring the measurements; this analyzes them. Pair it with your existing load tester (wrk, k6, JMH, etc.).
- A baseline manager. It compares two runs you give it; it doesn't store historical baselines or do drift-over-time tracking. (That's a natural extension — see the roadmap.)
- A profiler. It tells you that something regressed and by how much; it doesn't tell you where in the code. Use a profiler for that — the findings point you at which metric to profile.
- Magic for tiny samples. The distribution test declines (returns no finding) below 5 samples per side, because the normal approximation it uses isn't trustworthy there.
regression_radar/
stats.py Mann-Whitney U, Cliff's delta, percentiles — from scratch
model.py Run, MetricSeries, Finding, Comparison
detectors/ one file per detector
compare.py orchestrator: runs detectors, enforces evidence contract
loader.py JSON input
report.py Markdown + JSON renderers
llm.py optional LLM hook; no keys, no defaults
__main__.py CLI
tests/ 25 tests: stats properties, each detector, end-to-end
examples/ baseline + regressed + clean fixtures
docs/ STATISTICS.md, DESIGN.md
pip install pytest
pytest -v # 25 testsCI matrix runs on Python 3.10, 3.11, 3.12.
- Mann-Whitney U + Cliff's delta from scratch
- 5 detectors (latency dist, percentile, throughput, error rate, resource)
- Markdown + JSON reports,
--fail-on-regressionCI gate - Optional LLM narration hook (no keys baked in)
- 25 unit + end-to-end tests
- Historical baseline store — track drift across many runs, not just two
- Permutation test for small samples (where the normal approximation is weak)
- Direct adapters for wrk/k6/JMH output formats
- HTML report with distribution plots
- Hand off detected regressions to incident-pilot automatically
MIT.