Classify thousands of rows of free text — support tickets, reviews, complaints, inbound leads — with an LLM, and get a spreadsheet back with a category, a one-line summary, and a flag on the rows the model was unsure about. It reports its own accuracy and cost instead of asserting them.
All numbers are from an in-session run on 2026-09-06 with gemini-3.5-flash-lite
(Google AI Studio free tier), task complaint_routing, on a seeded sample of the
CFPB Consumer Complaint Database. $0.0000 is the free tier's price, not an
absence of calls — requests and tokens are shown alongside.
Full run — K = 20 (docs/eval.md):
| metric | value |
|---|---|
| accuracy | 0.804 |
| macro-F1 | 0.647 |
| rows scored | 2,540 of the committed 4,000 (see note) |
| requests | 171 |
| tokens | 644,684 (prompt 425,674 / completion 219,010) |
| measured cost | $0.0000 (Gemini free tier) |
| projected cost | ~$0.077 / 1,000 rows at OpenAI gpt-4o-mini list prices (2026-09-06), from the same measured token usage — a projection, not a paid run |
| errors | 0 |
| cache hit rate on an immediate re-run | 100% (0 requests) |
Partial run, by design-under-pressure. Around row 2,540 the free tier throttled from ~2 s to ~40 s per request. The run had checkpointed after every batch, so it stopped cleanly: 1,460 rows marked
pending(not dropped, not guessed), 0 errors, every figure here on the 2,540 scored rows, and one command —python scripts/full_run.py --rows 4000 --batch 20 --seed 20260906— resumes from row 2,541 without re-billing anything done. This is what the checkpoint / resume machinery is for.
Best classes: mortgage F1 0.96, student_loan 0.84. Weakest: personal_loan
(26 rows, F1 0.40) and other (5 rows, F1 0.00) — both nearly absent in the
source data. The most common confusion is debt_collection → credit_reporting
(109 of 747).
Batching trade-off — all three K run over the same 120 rows in one session, one task/prompt version, caches cleared first (docs/batching.md):
| K | accuracy | macro-F1 | requests / 1k rows (cold) | tokens / row |
|---|---|---|---|---|
| 1 | 0.825 | 0.633 | 1,000 | 906 |
| 10 | 0.842 | 0.655 | 100 | 276 |
| 20 | 0.842 | 0.649 | 50 | 231 |
No measurable accuracy penalty from batching here — K=20 landed +0.017 above K=1 (0.825 → 0.842), which at n = 120 is noise, not a real gain. Batching to K=20 also cuts tokens ~3.9× per row (906 → 231: the fixed system + label-definition
- few-shot scaffold is resent on every request, so it is paid once per row at K=1 and spread over 20 at K=20) and cuts requests 20×. On the free tier's 500 requests/day that is the difference between ~500 and ~10,000 rows classified in a day.
Label definitions vs terse names — 260 shared rows, K = 20 (docs/ablation.md): adding a one-line definition after each label name scored -0.023 accuracy / -0.020 macro-F1 versus bare names — a small effect, and in the opposite direction to what you'd assume. It edged out the batching delta (0.017); neither is large.
Review flag (docs/review_threshold.md): swept over
0.65 / 0.75 / 0.85 / 0.90. No threshold gives a clean trade-off. The flag is
precise — at 0.90 a flagged row is 2.3× more likely to be wrong than an
average row (0.46 vs a 0.20 base error rate) — but not sensitive: it still catches
only 26% of all errors while flagging 11% of rows. gemini-3.5-flash-lite is
confidently wrong often enough here that confidence alone is a weak error filter;
use the flag to prioritise a spot-check, not as a complete review queue. Measuring
this rather than assuming it is the point of the repo.
You have 5,000 support tickets in a spreadsheet. You want each one tagged with a category and a one-line summary, and you want the ones the model is not sure about marked so a human can look at just those. You also want to know, before you trust any of it, how accurate the tagging is against real labels and what it will cost per thousand rows.
A notebook that loops over rows calling an API does the first part for 200 rows. It falls over on 20,000: no resume after a crash, no cost ceiling, no signal on which rows are shaky, no accuracy number. This repo is the production version of that loop.
CSV ──▶ cache lookup ──▶ batch K rows/request ──▶ validate each row ──▶ CSV + XLSX
│ (sqlite) │ │
│ │ ├─ missing/invalid row → retry alone
│ │ └─ fails alone → marked error
▼ ▼
re-run = $0 checkpoint after every batch ──▶ --resume
- Batching — K rows per request (default 20). The free tier that makes this necessary allows 15 requests/minute and 500 requests/day: one request per row means 5,000 rows take ten days; 20 per request means one afternoon. Batching can also lower per-row accuracy, so the repo measures that (docs/batching.md) rather than assuming it.
- Row-by-row validation — the batch reply is checked per row against the label
list and a 0–1 confidence. A row missing from the reply or failing validation is
re-queued on its own; a row that fails alone is marked
errorwith the reason. - Cache — sqlite keyed by
sha256(model, task version, prompt version, row text). A re-run costs nothing and the hit rate is reported. - Resume — the orchestrator checkpoints to
work/after every batch.--resumecontinues; a run stopped by a budget or a daily quota marks the remaining rowspendingand exits cleanly. - Budget guard —
--budget 2.00raisesBudgetExhaustedbefore the call that would cross the ceiling. - Review flag — rows below the task's confidence threshold go on a
Reviewsheet with the reason. Whether that flag is actually informative is measured, not assumed (docs/eval.md).
No API key, no network:
pip install -e ".[dev]"
bulk-llm classify --demoThat runs a deterministic offline stub model over a bundled 200-row sample and
writes output/ (data.csv, review.csv, run.csv, results.xlsx). The stub
exercises the whole pipeline; it is not a real classifier.
To run a real model, copy .env.example to .env, fill in an OpenAI-compatible
endpoint, then:
bulk-llm classify --task complaint_routing --in data/sample.csv \
--text-col narrative --out output/ --batch 20 --budget 2.00Write a task YAML — no code change. Example (my_task.yaml):
name: churn_reason
version: "1"
description: Why did this customer cancel?
text_field: comment
review_threshold: 0.6
truncate_chars: 1000
labels:
- name: price
definition: The customer says it costs too much or a competitor is cheaper.
- name: missing_feature
definition: The customer needed something the product does not do.
- name: bug_or_reliability
definition: Crashes, downtime, data loss, or slow performance.
- name: other
definition: Anything else, or not enough information to tell.
few_shot:
- text: "Switched to a cheaper tool, could not justify the renewal."
label: price
summary: Renewed elsewhere for cost reasons.bulk-llm classify --task my_task.yaml --in cancellations.csv --text-col comment --out output/The shipped tasks are complaint_routing, sentiment, and lead_qualification
(src/bulk_llm/tasks/).
The sample is 4,000 rows of the CFPB Consumer Complaint Database (US
government, public domain). CFPB redacts personal data to XXXX before
publishing narratives, which is why the sample is safe to commit.
scripts/check_no_pii.py scans every tracked file for emails, phone numbers,
SSNs and long digit runs and fails CI on a hit. Full source, licence, sampling
command, seed and the Product → label mapping: data/README.md.
- Accuracy depends on how the labels are defined. Here, adding one-line definitions moved accuracy by -0.023 versus bare label names — small, and the opposite of the intuitive direction (docs/ablation.md). On another dataset it could be larger, or flip. Measure it for yours.
- Batching can lower per-row accuracy — but didn't measurably here. On the same 120 rows K=20 scored +0.017 above K=1 (docs/batching.md) — noise at that sample size, so read it as "no measurable penalty", not a gain. On another dataset or model the sign could go the other way; measure it.
- The labels are what a consumer selected, not a gold annotation. The 0.804
accuracy is agreement with the CFPB
Productfield — a proxy for correct routing, not ground truth. TheConsumer Loan→vehicle_loanmerge in the mapping adds its own noise. - The confidence flag is a weak error filter on this model. Swept 0.65–0.90 (docs/review_threshold.md): flagged rows are ~2.3× likelier to be wrong than average, but even at 0.90 the flag catches only 26% of errors. Good for a spot-check, not a full review queue. Measured, not hidden.
- The headline run is on 2,540 of 4,000 rows. The free tier throttled mid-run;
the rest are
pendingwith a resume command in docs/eval.md. - Cost figures are for one model on one date —
gemini-3.5-flash-lite, free tier, 2026-09-06. A different model or a paid tier will differ; re-run the scripts for your own numbers. - Committed narratives are truncated to 500 characters to keep the sample under 2 MB, so these rows cap what any model here can see.
Python 3.10+, httpx, PyYAML, openpyxl, tenacity; ruff + pytest for
dev; matplotlib for the charts. src/ layout, GitHub Actions CI on 3.10 and
3.12.
MIT — see LICENSE.

