Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

groundedtune

CI License: MIT Python 3.9+

Safe, research-grounded continual fine-tuning for local LLMs — learns from real external sources and your corrections, never from its own unchecked output, with regression checks and human approval before every update.

Твоя локальная модель, которая реально учится — безопасно, с доказательствами.

The problem: model collapse

Naively fine-tuning a model on its own unchecked outputs is a documented failure mode — "model collapse." Each generation trains on the previous generation's mistakes and drift, diversity narrows, and errors amplify with every cycle. This is well studied in the literature on recursive training on synthetic data — see Shumailov et al., "AI models collapse when trained on recursively generated data", Nature (2024).

There is no ready-made, safe, continual-fine-tuning tool for local models on consumer hardware:

  • Axolotl / LlamaFactory / plain TRL scripts give you a solid training loop, but assume you already have a clean, provenance-tagged dataset — they have no opinion on where your data came from or whether it's trustworthy.
  • torchtune is a great low-level recipe library for the training step itself, with the same assumption: bring your own vetted dataset.
  • groundedtune sits one layer earlier: it produces that vetted dataset, and structurally refuses to let ungrounded model output into it. M2 hands the resulting dataset to Unsloth/TRL for the actual training loop (see Roadmap) rather than reinventing it.

groundedtune's answer: every example that goes into a training dataset carries provenance back to something that is not the model's own unverified output — an external source a human supplied, or a correction a human made. If groundedtune can't trace an example to one of those, it doesn't go in the dataset.

Try it in 30 seconds

No setup, no files to author yourself — examples/ has everything committed:

git clone https://github.com/tsvirov/groundedtune.git && cd groundedtune
python3 -m venv .venv && .venv/bin/pip install -e . -q
source .venv/bin/activate
bash examples/walkthrough.sh

It ingests a human correction, a human-supplied claim + grounding text, and an assumecheck decision log into one dataset, then prints dataset stats and exports the SFT/DPO rows. See examples/README.md for the real, captured output.

Prefer a narrated version? pip install rich && python3 examples/hero_demo.py — the same real sample log, staged for a GIF (clearly labeled as scripted; the command above is the actual CLI):

groundedtune hero demo

Scope of this build

This repository implements all four planned milestones:

  • M1 — the data pipeline with provenance. Pydantic v2 schemas for provenance-tagged training examples (SFTExample, DPOExample); source adapters that turn a trustworthy signal into a tagged example (human corrections, human-supplied manual grounding, and assumecheck decision logs); an append-only, dedup-on-write JSONL dataset store; a groundedtune CLI to ingest sources and inspect/export the resulting dataset.
  • M2 — LoRA/DPO training orchestration, M3 — regression eval harness with diff report, and M4 — CLI versioning/rollback — see Roadmap below for what each one does, and for the one honest caveat in M2 (UnslothTrainerBackend is real code that has never been run end to end on this project's dev machine, which has no NVIDIA GPU).

Architecture

sources/                 -> examples with provenance
  correction.py             human draft -> final rewrite      (DPOExample)
  manual.py                 human claim + grounding text       (SFTExample)
  assumecheck_log.py        assumecheck's FLAG/PASS log        (DPOExample | SFTExample)
        |
        v
  schema.py               ProvenanceRecord, SFTExample, DPOExample (Pydantic v2)
        |
        v
  store.py                DatasetStore: append-only JSONL, dedup by content_hash
        |
        v
  cli.py                  `groundedtune ingest ...` / `groundedtune dataset ...`

Every example's ProvenanceRecord records:

  • source_type — one of web_research, user_correction, assumecheck_log, manual
  • source_ref — a URL, file path, or log line reference identifying the origin
  • confidence — how much this example should be trusted (0..1)
  • timestamp — when the example entered the pipeline (ISO8601, always generated at ingestion time)
  • content_hash — sha256(prompt + "\n" + response_or_chosen), used for dataset dedup

SFTExample.prompt/response and DPOExample.prompt/chosen/rejected must all be non-empty (whitespace-only doesn't count either) — schema validation rejects a content-free "grounded" example before it can ever be written to a dataset, regardless of which source adapter or CLI command tried to build it.

Source adapters

sources/correction.py — from_correction(prompt, draft, final) Builds a DPOExample from a human correction: chosen=final, rejected=draft. A direct human correction is the most trustworthy signal groundedtune can ingest, so confidence=1.0 and there is no external source_ref — the human is the source. Raises ValueError if draft and final are identical (after stripping) — that's not a correction, and would otherwise silently produce a degenerate chosen == rejected DPO pair at the highest possible confidence.

sources/manual.py — from_manual_source(claim, source_text, source_url, confidence=0.8) For when a human manually supplies a claim plus the source text that grounds it. M1 intentionally has no automated web fetching — nothing ungrounded should slip in without a human looking at it first. Automated web-search -> fetch -> verify is planned for a later milestone (source_type="web_research"), reusing the same search -> fetch -> verify pattern used by internal deep-research workflows; until then, a human supplies the grounding text directly.

sources/assumecheck_log.py — from_assumecheck_log(log_path) Parses assumecheck's JSONL decision log and returns (examples, stats):

assumecheck line groundedtune result
decision == "FLAG" and resolution present DPOExample: chosen = the resolution (combined with the consensus goal, if present), rejected = the first divergent sample serialized as text. confidence=0.9 — a real, externally-resolved disagreement.
decision == "FLAG" and resolution missing Skipped, but counted in stats["pending_resolution"] — not silently dropped, just not trainable yet.
decision == "PASS" with a consensus object Optional low-confidence SFTExample built from the consensus goal + assumptions. confidence=0.5 — this is a weaker signal than the others: the samples merely agreed with themselves, which is not the same as being externally verified.
malformed JSON line Skipped with a stderr warning, counted in stats["malformed"]; ingestion never aborts because one line in a long log is broken.
a row that parses as JSON but has the wrong shape for the branch it hits (e.g. consensus is a string instead of an object, samples is a string instead of a list, or a required field like prompt is missing/blank) Skipped with a stderr warning, counted in stats["skipped_bad_shape"] — never silently manufactures a content-free example, and never aborts the rest of the log.

Dataset format

A dataset is a single JSONL file. Every line is one example, tagged with "kind": "sft" or "kind": "dpo" so a mixed file round-trips correctly:

{"kind": "sft", "prompt": "What year was it founded?", "response": "1889.", "provenance": {"source_type": "manual", "source_ref": "https://example.com", "confidence": 0.8, "timestamp": "2026-01-01T00:00:00+00:00", "content_hash": "..."}}
{"kind": "dpo", "prompt": "Rewrite this.", "chosen": "good text", "rejected": "bad text", "provenance": {"source_type": "user_correction", "source_ref": null, "confidence": 1.0, "timestamp": "2026-01-01T00:00:00+00:00", "content_hash": "..."}}

DatasetStore.append() dedups by content_hash: appending an example whose hash already exists in the file is a no-op (it returns False). A malformed or schema-invalid line encountered while reading a dataset (e.g. from a hand-edited file, or a crash mid-write) is skipped with a stderr warning rather than crashing dataset stats/dataset export.

Two things worth knowing about the dedup key specifically, since the formula is fixed by the M1 spec (sha256(prompt + "\n" + response_or_chosen)):

  • It is not injective across the prompt/response boundary — distinct (prompt, response) pairs whose concatenated prompt + "\n" + response strings happen to be identical will hash identically and collide.
  • For DPOExample, only prompt and chosen participate in the hash — rejected is not part of the dedup key. Two DPO pairs with the same (prompt, chosen) but a genuinely different rejected will collide, and only the first one ingested is kept.

Install

Not yet on PyPI. Install from source:

git clone https://github.com/tsvirov/groundedtune.git
cd groundedtune
pip install -e .

Requires Python >= 3.9.

CLI usage

# A human correction (draft -> final) becomes a DPO pair.
groundedtune ingest correction \
  --draft examples/draft.txt --final examples/final.txt \
  --prompt "What year was the library founded?" \
  --out dataset.jsonl

# A human-supplied claim + grounding text becomes an SFT example.
# --source-text is always literal text; --source-text-file reads a file.
# Exactly one of the two is required — there is no auto-detection between
# them, so a literal value is never silently reinterpreted as a file path.
groundedtune ingest manual \
  --claim "What year was the library founded?" \
  --source-text-file examples/source.txt \
  --source-url https://example.com/history \
  --confidence 0.8 \
  --out dataset.jsonl

# Ingest an assumecheck decision log.
groundedtune ingest assumecheck-log examples/sample-assumecheck-log.jsonl --out dataset.jsonl
# -> prints how many examples were added, and how many are pending_resolution.

# Inspect a dataset.
groundedtune dataset stats dataset.jsonl

# Export just the SFT or DPO rows.
groundedtune dataset export dataset.jsonl --format sft --out sft.jsonl
groundedtune dataset export dataset.jsonl --format dpo --out dpo.jsonl

# M2: train a LoRA adapter from a dataset ("mock" is deterministic/GPU-free;
# "unsloth" needs a real CUDA GPU — see Roadmap).
groundedtune train --dataset dataset.jsonl --base-model qwen3.5-9b \
  --backend mock --output-dir adapter-v1

# M3: diff a before/after backend pair against a fixed eval set.
# Exit code 1 if anything newly fails — this one is a CI gate.
groundedtune eval --before-backend mock --before-config before.json \
  --after-backend mock --after-config after.json \
  --eval-set evalsets/core-v1.jsonl --out report.json

# M4: record a version (starts unapproved), approve it, then point 'current'
# at it. set-current/rollback refuse to target an unapproved version.
groundedtune version record --dataset dataset.jsonl --adapter-dir adapter-v1 \
  --eval-report report.json
groundedtune version approve 1
groundedtune version set-current 1
groundedtune version list
groundedtune version show 1
groundedtune version rollback

groundedtune --version

Validation failures (an out-of-range --confidence, a degenerate/blank example, a malformed dataset row) are reported as a clean, single-line CLI error — never a raw Python traceback. See examples/m2-m4-walkthrough.sh for a full, real, captured run of the M2-M4 commands above.

Integration with assumecheck

groundedtune consumes assumecheck's log.jsonl — the record of every time assumecheck ran multiple samples against a prompt, detected divergence above threshold, and either flagged it for a human or passed it because the samples agreed. Each line's samples field is a list of assumecheck's IntentExtraction objects (goal, given_inputs, missing_inputs, assumptions, ambiguity_flags); consensus is null on FLAG lines and one such object on PASS lines. The field groundedtune cares about most is resolution: assumecheck itself always logs it as null — it's a forward-compatible slot a downstream tool or human-in-the-loop wrapper fills in later. Once a human resolves a FLAG, that resolution is a genuine, externally-verified correction and becomes training data with confidence=0.9. Until a FLAG line has a resolution, groundedtune leaves it out of the dataset but still counts it, so you always know how much signal is sitting there waiting for a human.

See examples/sample-assumecheck-log.jsonl for a realistic log exercising every branch.

Limitations (M1)

Known characteristics of the current build, not roadmap gaps:

  • Exact-hash dedup only. content_hash is an exact match on sha256(prompt + "\n" + response_or_chosen) — there is no near-duplicate or fuzzy-match detection. Two examples that say the same thing in different words are both kept.
  • Single flat JSONL file, no merge/split tooling. A dataset is one file; combining or splitting datasets across files is a manual cat/split operation, not something the CLI does for you.
  • source_ref is not verified reachable. It's a free-text URL/path/log-line reference for human traceability — groundedtune doesn't check that it still resolves to anything.
  • Confidence values are fixed constants per source adapter, not learned or recalibrated from outcomes (1.0 for corrections, 0.8 default for manual, 0.9/0.5 for assumecheck FLAG/PASS).

Roadmap

M2, M3, and M4 are now implemented. Read this section precisely — one part of M2 is real code that has never been run end to end, and that's stated here on purpose rather than glossed over.

  • M2 — LoRA/DPO training orchestration. groundedtune train --dataset ... --base-model ... --backend mock|unsloth --output-dir ... (src/groundedtune/train/backend.py). A TrainerBackend protocol with two implementations:
    • MockTrainerBackend — deterministic, network- and GPU-free. Reuses DatasetStore.load to read/validate the dataset, writes a fake adapter directory (adapter_config.json + a marker file) and a metrics.json with a fake loss curve. This is the only backend exercised in this project's tests, CLI demo, and CI, on every platform.
    • UnslothTrainerBackend — real Unsloth/TRL/PEFT wiring (lazy imports, FastLanguageModel + LoRA + SFTTrainer/DPOTrainer). It requires an NVIDIA GPU with CUDA; if imports fail or no CUDA device is found, it raises a clear RuntimeError with install/hardware instructions — it never silently falls back to the mock and never pretends to have trained anything. Honesty note: this repository was built and tested on a Mac with no NVIDIA GPU. Only the import/CUDA guard is tested (tests/test_train_backend.py, via a monkeypatched failing import) — the training loop itself (model loading, PEFT wrapping, SFTTrainer/DPOTrainer execution, adapter save) has never been run end to end anywhere. Treat it as reviewed-but-unverified code, not a tested path, until someone runs it on real CUDA hardware.
  • M3 — Regression eval harness with diff report. groundedtune eval --before-backend mock|ollama --before-config before.json --after-backend mock|ollama --after-config after.json --eval-set evalsets/core-v1.jsonl [--out report.json] (src/groundedtune/eval/). An EvalBackend protocol (MockEvalBackend for canned prompt->response maps, OllamaEvalBackend for a real local Ollama-compatible /api/chat server, mirroring assumecheck's runtime.py pattern — tested only via httpx.MockTransport, no live server needed). The eval set format is a deliberately simple JSONL of {"prompt", "check": {"type": "contains"|"not_contains"|"regex", "value"}} — no LLM-judge dependency. harness.py runs a before/after backend pair against the same eval set and buckets results into newly_passing / newly_failing / stayed_passing / stayed_failing. Unlike assumecheck's eval command, this one is a CI gate: exit code 0 if newly_failing is empty, exit code 1 if there are regressions.
  • M4 — CLI versioning/rollback. groundedtune version record|approve|list|show|set-current|rollback (src/groundedtune/versioning.py). A manifest at .groundedtune/versions.json tracks each recorded version (dataset path + sha256 hash, adapter dir, eval report summary, timestamp, approved: bool, default false). Critical, tested invariant: set-current and rollback refuse — clean error, non-zero exit — to point current at a version that isn't approved: true (tests/test_versioning.py::test_set_current_on_unapproved_version_is_refused and the CLI-level equivalent in tests/test_cli_version.py). rollback sets current to the previous approved version before the current one, skipping any unapproved versions in between. This MVP only tracks and gates the pointer — it does not deploy or serve anything.

See examples/m2-m4-walkthrough.sh and examples/README.md for a real, captured run of train -> eval -> version record -> approve -> set-current, including the two steps that are supposed to fail (an unapproved set-current, a rollback with nothing earlier to roll back to).

Development

python3 -m venv .venv
.venv/bin/pip install -U pip
.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q
.venv/bin/ruff check .

License

MIT — see LICENSE.

About

Safe, research-grounded continual fine-tuning for local LLMs — learns from real external sources and your corrections, never from its own unchecked output, with regression checks and human approval before every update.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages