Skip to content

Repository files navigation

llm-bulk-classifier

Classify thousands of rows of free text — support tickets, reviews, complaints, inbound leads — with an LLM, and get a spreadsheet back with a category, a one-line summary, and a flag on the rows the model was unsure about. It reports its own accuracy and cost instead of asserting them.

Data sheet preview

Results

All numbers are from an in-session run on 2026-09-06 with gemini-3.5-flash-lite (Google AI Studio free tier), task complaint_routing, on a seeded sample of the CFPB Consumer Complaint Database. $0.0000 is the free tier's price, not an absence of calls — requests and tokens are shown alongside.

Full run — K = 20 (docs/eval.md):

metric value
accuracy 0.804
macro-F1 0.647
rows scored 2,540 of the committed 4,000 (see note)
requests 171
tokens 644,684 (prompt 425,674 / completion 219,010)
measured cost $0.0000 (Gemini free tier)
projected cost ~$0.077 / 1,000 rows at OpenAI gpt-4o-mini list prices (2026-09-06), from the same measured token usage — a projection, not a paid run
errors 0
cache hit rate on an immediate re-run 100% (0 requests)

Partial run, by design-under-pressure. Around row 2,540 the free tier throttled from ~2 s to ~40 s per request. The run had checkpointed after every batch, so it stopped cleanly: 1,460 rows marked pending (not dropped, not guessed), 0 errors, every figure here on the 2,540 scored rows, and one command — python scripts/full_run.py --rows 4000 --batch 20 --seed 20260906 — resumes from row 2,541 without re-billing anything done. This is what the checkpoint / resume machinery is for.

Best classes: mortgage F1 0.96, student_loan 0.84. Weakest: personal_loan (26 rows, F1 0.40) and other (5 rows, F1 0.00) — both nearly absent in the source data. The most common confusion is debt_collection → credit_reporting (109 of 747).

Batching trade-off — all three K run over the same 120 rows in one session, one task/prompt version, caches cleared first (docs/batching.md):

K accuracy macro-F1 requests / 1k rows (cold) tokens / row
1 0.825 0.633 1,000 906
10 0.842 0.655 100 276
20 0.842 0.649 50 231

No measurable accuracy penalty from batching here — K=20 landed +0.017 above K=1 (0.825 → 0.842), which at n = 120 is noise, not a real gain. Batching to K=20 also cuts tokens ~3.9× per row (906 → 231: the fixed system + label-definition

  • few-shot scaffold is resent on every request, so it is paid once per row at K=1 and spread over 20 at K=20) and cuts requests 20×. On the free tier's 500 requests/day that is the difference between ~500 and ~10,000 rows classified in a day.

Batching curve

Label definitions vs terse names — 260 shared rows, K = 20 (docs/ablation.md): adding a one-line definition after each label name scored -0.023 accuracy / -0.020 macro-F1 versus bare names — a small effect, and in the opposite direction to what you'd assume. It edged out the batching delta (0.017); neither is large.

Review flag (docs/review_threshold.md): swept over 0.65 / 0.75 / 0.85 / 0.90. No threshold gives a clean trade-off. The flag is precise — at 0.90 a flagged row is 2.3× more likely to be wrong than an average row (0.46 vs a 0.20 base error rate) — but not sensitive: it still catches only 26% of all errors while flagging 11% of rows. gemini-3.5-flash-lite is confidently wrong often enough here that confidence alone is a weak error filter; use the flag to prioritise a spot-check, not as a complete review queue. Measuring this rather than assuming it is the point of the repo.

The job this solves

You have 5,000 support tickets in a spreadsheet. You want each one tagged with a category and a one-line summary, and you want the ones the model is not sure about marked so a human can look at just those. You also want to know, before you trust any of it, how accurate the tagging is against real labels and what it will cost per thousand rows.

A notebook that loops over rows calling an API does the first part for 200 rows. It falls over on 20,000: no resume after a crash, no cost ceiling, no signal on which rows are shaky, no accuracy number. This repo is the production version of that loop.

How it works

CSV ──▶ cache lookup ──▶ batch K rows/request ──▶ validate each row ──▶ CSV + XLSX
         │ (sqlite)          │                       │
         │                   │                       ├─ missing/invalid row → retry alone
         │                   │                       └─ fails alone → marked error
         ▼                   ▼
     re-run = $0        checkpoint after every batch  ──▶  --resume
  • Batching — K rows per request (default 20). The free tier that makes this necessary allows 15 requests/minute and 500 requests/day: one request per row means 5,000 rows take ten days; 20 per request means one afternoon. Batching can also lower per-row accuracy, so the repo measures that (docs/batching.md) rather than assuming it.
  • Row-by-row validation — the batch reply is checked per row against the label list and a 0–1 confidence. A row missing from the reply or failing validation is re-queued on its own; a row that fails alone is marked error with the reason.
  • Cache — sqlite keyed by sha256(model, task version, prompt version, row text). A re-run costs nothing and the hit rate is reported.
  • Resume — the orchestrator checkpoints to work/ after every batch. --resume continues; a run stopped by a budget or a daily quota marks the remaining rows pending and exits cleanly.
  • Budget guard — --budget 2.00 raises BudgetExhausted before the call that would cross the ceiling.
  • Review flag — rows below the task's confidence threshold go on a Review sheet with the reason. Whether that flag is actually informative is measured, not assumed (docs/eval.md).

Quick start

No API key, no network:

pip install -e ".[dev]"
bulk-llm classify --demo

That runs a deterministic offline stub model over a bundled 200-row sample and writes output/ (data.csv, review.csv, run.csv, results.xlsx). The stub exercises the whole pipeline; it is not a real classifier.

To run a real model, copy .env.example to .env, fill in an OpenAI-compatible endpoint, then:

bulk-llm classify --task complaint_routing --in data/sample.csv \
    --text-col narrative --out output/ --batch 20 --budget 2.00

Adapting to your data

Write a task YAML — no code change. Example (my_task.yaml):

name: churn_reason
version: "1"
description: Why did this customer cancel?
text_field: comment
review_threshold: 0.6
truncate_chars: 1000
labels:
  - name: price
    definition: The customer says it costs too much or a competitor is cheaper.
  - name: missing_feature
    definition: The customer needed something the product does not do.
  - name: bug_or_reliability
    definition: Crashes, downtime, data loss, or slow performance.
  - name: other
    definition: Anything else, or not enough information to tell.
few_shot:
  - text: "Switched to a cheaper tool, could not justify the renewal."
    label: price
    summary: Renewed elsewhere for cost reasons.
bulk-llm classify --task my_task.yaml --in cancellations.csv --text-col comment --out output/

The shipped tasks are complaint_routing, sentiment, and lead_qualification (src/bulk_llm/tasks/).

Data & privacy

The sample is 4,000 rows of the CFPB Consumer Complaint Database (US government, public domain). CFPB redacts personal data to XXXX before publishing narratives, which is why the sample is safe to commit. scripts/check_no_pii.py scans every tracked file for emails, phone numbers, SSNs and long digit runs and fails CI on a hit. Full source, licence, sampling command, seed and the Product → label mapping: data/README.md.

Limitations

  • Accuracy depends on how the labels are defined. Here, adding one-line definitions moved accuracy by -0.023 versus bare label names — small, and the opposite of the intuitive direction (docs/ablation.md). On another dataset it could be larger, or flip. Measure it for yours.
  • Batching can lower per-row accuracy — but didn't measurably here. On the same 120 rows K=20 scored +0.017 above K=1 (docs/batching.md) — noise at that sample size, so read it as "no measurable penalty", not a gain. On another dataset or model the sign could go the other way; measure it.
  • The labels are what a consumer selected, not a gold annotation. The 0.804 accuracy is agreement with the CFPB Product field — a proxy for correct routing, not ground truth. The Consumer Loan → vehicle_loan merge in the mapping adds its own noise.
  • The confidence flag is a weak error filter on this model. Swept 0.65–0.90 (docs/review_threshold.md): flagged rows are ~2.3× likelier to be wrong than average, but even at 0.90 the flag catches only 26% of errors. Good for a spot-check, not a full review queue. Measured, not hidden.
  • The headline run is on 2,540 of 4,000 rows. The free tier throttled mid-run; the rest are pending with a resume command in docs/eval.md.
  • Cost figures are for one model on one date — gemini-3.5-flash-lite, free tier, 2026-09-06. A different model or a paid tier will differ; re-run the scripts for your own numbers.
  • Committed narratives are truncated to 500 characters to keep the sample under 2 MB, so these rows cap what any model here can see.

Stack

Python 3.10+, httpx, PyYAML, openpyxl, tenacity; ruff + pytest for dev; matplotlib for the charts. src/ layout, GitHub Actions CI on 3.10 and 3.12.

License

MIT — see LICENSE.

About

Classify thousands of rows of free text with an LLM: batching, cache, resume, USD budget guard, low-confidence review flags, and accuracy measured against real labels. Reports its own cost and accuracy instead of asserting them.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages