Safe, research-grounded continual fine-tuning for local LLMs — learns from real external sources and your corrections, never from its own unchecked output, with regression checks and human approval before every update.
Твоя локальная модель, которая реально учится — безопасно, с доказательствами.
Naively fine-tuning a model on its own unchecked outputs is a documented failure mode — "model collapse." Each generation trains on the previous generation's mistakes and drift, diversity narrows, and errors amplify with every cycle. This is well studied in the literature on recursive training on synthetic data — see Shumailov et al., "AI models collapse when trained on recursively generated data", Nature (2024).
There is no ready-made, safe, continual-fine-tuning tool for local models on consumer hardware:
- Axolotl / LlamaFactory / plain TRL scripts give you a solid training loop, but assume you already have a clean, provenance-tagged dataset — they have no opinion on where your data came from or whether it's trustworthy.
- torchtune is a great low-level recipe library for the training step itself, with the same assumption: bring your own vetted dataset.
- groundedtune sits one layer earlier: it produces that vetted dataset, and structurally refuses to let ungrounded model output into it. M2 hands the resulting dataset to Unsloth/TRL for the actual training loop (see Roadmap) rather than reinventing it.
groundedtune's answer: every example that goes into a training dataset carries provenance back to something that is not the model's own unverified output — an external source a human supplied, or a correction a human made. If groundedtune can't trace an example to one of those, it doesn't go in the dataset.
No setup, no files to author yourself — examples/ has
everything committed:
git clone https://github.com/tsvirov/groundedtune.git && cd groundedtune
python3 -m venv .venv && .venv/bin/pip install -e . -q
source .venv/bin/activate
bash examples/walkthrough.shIt ingests a human correction, a human-supplied claim + grounding text, and
an assumecheck decision log into one dataset, then prints dataset stats
and exports the SFT/DPO rows. See examples/README.md
for the real, captured output.
Prefer a narrated version? pip install rich && python3 examples/hero_demo.py — the same real sample log, staged for a GIF (clearly
labeled as scripted; the command above is the actual CLI):
This repository implements all four planned milestones:
- M1 — the data pipeline with provenance. Pydantic v2 schemas for
provenance-tagged training examples (
SFTExample,DPOExample); source adapters that turn a trustworthy signal into a tagged example (human corrections, human-supplied manual grounding, and assumecheck decision logs); an append-only, dedup-on-write JSONL dataset store; agroundedtuneCLI to ingest sources and inspect/export the resulting dataset. - M2 — LoRA/DPO training orchestration, M3 — regression eval harness
with diff report, and M4 — CLI versioning/rollback — see
Roadmap below for what each one does, and for the one honest
caveat in M2 (
UnslothTrainerBackendis real code that has never been run end to end on this project's dev machine, which has no NVIDIA GPU).
sources/ -> examples with provenance
correction.py human draft -> final rewrite (DPOExample)
manual.py human claim + grounding text (SFTExample)
assumecheck_log.py assumecheck's FLAG/PASS log (DPOExample | SFTExample)
|
v
schema.py ProvenanceRecord, SFTExample, DPOExample (Pydantic v2)
|
v
store.py DatasetStore: append-only JSONL, dedup by content_hash
|
v
cli.py `groundedtune ingest ...` / `groundedtune dataset ...`
Every example's ProvenanceRecord records:
source_type— one ofweb_research,user_correction,assumecheck_log,manualsource_ref— a URL, file path, or log line reference identifying the originconfidence— how much this example should be trusted (0..1)timestamp— when the example entered the pipeline (ISO8601, always generated at ingestion time)content_hash—sha256(prompt + "\n" + response_or_chosen), used for dataset dedup
SFTExample.prompt/response and DPOExample.prompt/chosen/rejected
must all be non-empty (whitespace-only doesn't count either) — schema
validation rejects a content-free "grounded" example before it can ever be
written to a dataset, regardless of which source adapter or CLI command
tried to build it.
sources/correction.py — from_correction(prompt, draft, final)
Builds a DPOExample from a human correction: chosen=final,
rejected=draft. A direct human correction is the most trustworthy signal
groundedtune can ingest, so confidence=1.0 and there is no external
source_ref — the human is the source. Raises ValueError if draft and
final are identical (after stripping) — that's not a correction, and would
otherwise silently produce a degenerate chosen == rejected DPO pair at the
highest possible confidence.
sources/manual.py — from_manual_source(claim, source_text, source_url, confidence=0.8)
For when a human manually supplies a claim plus the source text that grounds
it. M1 intentionally has no automated web fetching — nothing ungrounded
should slip in without a human looking at it first. Automated
web-search -> fetch -> verify is planned for a later milestone
(source_type="web_research"), reusing the same search -> fetch -> verify
pattern used by internal deep-research workflows; until then, a human
supplies the grounding text directly.
sources/assumecheck_log.py — from_assumecheck_log(log_path)
Parses assumecheck's JSONL decision log and
returns (examples, stats):
| assumecheck line | groundedtune result |
|---|---|
decision == "FLAG" and resolution present |
DPOExample: chosen = the resolution (combined with the consensus goal, if present), rejected = the first divergent sample serialized as text. confidence=0.9 — a real, externally-resolved disagreement. |
decision == "FLAG" and resolution missing |
Skipped, but counted in stats["pending_resolution"] — not silently dropped, just not trainable yet. |
decision == "PASS" with a consensus object |
Optional low-confidence SFTExample built from the consensus goal + assumptions. confidence=0.5 — this is a weaker signal than the others: the samples merely agreed with themselves, which is not the same as being externally verified. |
| malformed JSON line | Skipped with a stderr warning, counted in stats["malformed"]; ingestion never aborts because one line in a long log is broken. |
a row that parses as JSON but has the wrong shape for the branch it hits (e.g. consensus is a string instead of an object, samples is a string instead of a list, or a required field like prompt is missing/blank) |
Skipped with a stderr warning, counted in stats["skipped_bad_shape"] — never silently manufactures a content-free example, and never aborts the rest of the log. |
A dataset is a single JSONL file. Every line is one example, tagged with
"kind": "sft" or "kind": "dpo" so a mixed file round-trips correctly:
{"kind": "sft", "prompt": "What year was it founded?", "response": "1889.", "provenance": {"source_type": "manual", "source_ref": "https://example.com", "confidence": 0.8, "timestamp": "2026-01-01T00:00:00+00:00", "content_hash": "..."}}
{"kind": "dpo", "prompt": "Rewrite this.", "chosen": "good text", "rejected": "bad text", "provenance": {"source_type": "user_correction", "source_ref": null, "confidence": 1.0, "timestamp": "2026-01-01T00:00:00+00:00", "content_hash": "..."}}DatasetStore.append() dedups by content_hash: appending an example whose
hash already exists in the file is a no-op (it returns False). A malformed
or schema-invalid line encountered while reading a dataset (e.g. from a
hand-edited file, or a crash mid-write) is skipped with a stderr warning
rather than crashing dataset stats/dataset export.
Two things worth knowing about the dedup key specifically, since the formula
is fixed by the M1 spec (sha256(prompt + "\n" + response_or_chosen)):
- It is not injective across the prompt/response boundary — distinct
(prompt, response)pairs whose concatenatedprompt + "\n" + responsestrings happen to be identical will hash identically and collide. - For
DPOExample, onlypromptandchosenparticipate in the hash —rejectedis not part of the dedup key. Two DPO pairs with the same(prompt, chosen)but a genuinely differentrejectedwill collide, and only the first one ingested is kept.
Not yet on PyPI. Install from source:
git clone https://github.com/tsvirov/groundedtune.git
cd groundedtune
pip install -e .Requires Python >= 3.9.
# A human correction (draft -> final) becomes a DPO pair.
groundedtune ingest correction \
--draft examples/draft.txt --final examples/final.txt \
--prompt "What year was the library founded?" \
--out dataset.jsonl
# A human-supplied claim + grounding text becomes an SFT example.
# --source-text is always literal text; --source-text-file reads a file.
# Exactly one of the two is required — there is no auto-detection between
# them, so a literal value is never silently reinterpreted as a file path.
groundedtune ingest manual \
--claim "What year was the library founded?" \
--source-text-file examples/source.txt \
--source-url https://example.com/history \
--confidence 0.8 \
--out dataset.jsonl
# Ingest an assumecheck decision log.
groundedtune ingest assumecheck-log examples/sample-assumecheck-log.jsonl --out dataset.jsonl
# -> prints how many examples were added, and how many are pending_resolution.
# Inspect a dataset.
groundedtune dataset stats dataset.jsonl
# Export just the SFT or DPO rows.
groundedtune dataset export dataset.jsonl --format sft --out sft.jsonl
groundedtune dataset export dataset.jsonl --format dpo --out dpo.jsonl
# M2: train a LoRA adapter from a dataset ("mock" is deterministic/GPU-free;
# "unsloth" needs a real CUDA GPU — see Roadmap).
groundedtune train --dataset dataset.jsonl --base-model qwen3.5-9b \
--backend mock --output-dir adapter-v1
# M3: diff a before/after backend pair against a fixed eval set.
# Exit code 1 if anything newly fails — this one is a CI gate.
groundedtune eval --before-backend mock --before-config before.json \
--after-backend mock --after-config after.json \
--eval-set evalsets/core-v1.jsonl --out report.json
# M4: record a version (starts unapproved), approve it, then point 'current'
# at it. set-current/rollback refuse to target an unapproved version.
groundedtune version record --dataset dataset.jsonl --adapter-dir adapter-v1 \
--eval-report report.json
groundedtune version approve 1
groundedtune version set-current 1
groundedtune version list
groundedtune version show 1
groundedtune version rollback
groundedtune --versionValidation failures (an out-of-range --confidence, a degenerate/blank
example, a malformed dataset row) are reported as a clean, single-line CLI
error — never a raw Python traceback. See
examples/m2-m4-walkthrough.sh for a full,
real, captured run of the M2-M4 commands above.
groundedtune consumes assumecheck's log.jsonl —
the record of every time assumecheck ran multiple samples against a prompt,
detected divergence above threshold, and either flagged it for a human or
passed it because the samples agreed. Each line's samples field is a list
of assumecheck's IntentExtraction objects (goal, given_inputs,
missing_inputs, assumptions, ambiguity_flags); consensus is null on
FLAG lines and one such object on PASS lines. The field groundedtune
cares about most is resolution: assumecheck itself always logs it as
null — it's a forward-compatible slot a downstream tool or
human-in-the-loop wrapper fills in later. Once a human resolves a FLAG,
that resolution is a genuine, externally-verified correction and becomes
training data with confidence=0.9. Until a FLAG line has a resolution,
groundedtune leaves it out of the dataset but still counts it, so you always
know how much signal is sitting there waiting for a human.
See examples/sample-assumecheck-log.jsonl
for a realistic log exercising every branch.
Known characteristics of the current build, not roadmap gaps:
- Exact-hash dedup only.
content_hashis an exact match onsha256(prompt + "\n" + response_or_chosen)— there is no near-duplicate or fuzzy-match detection. Two examples that say the same thing in different words are both kept. - Single flat JSONL file, no merge/split tooling. A dataset is one file;
combining or splitting datasets across files is a manual
cat/splitoperation, not something the CLI does for you. source_refis not verified reachable. It's a free-text URL/path/log-line reference for human traceability — groundedtune doesn't check that it still resolves to anything.- Confidence values are fixed constants per source adapter, not learned
or recalibrated from outcomes (
1.0for corrections,0.8default for manual,0.9/0.5for assumecheck FLAG/PASS).
M2, M3, and M4 are now implemented. Read this section precisely — one part of M2 is real code that has never been run end to end, and that's stated here on purpose rather than glossed over.
- M2 — LoRA/DPO training orchestration.
groundedtune train --dataset ... --base-model ... --backend mock|unsloth --output-dir ...(src/groundedtune/train/backend.py). ATrainerBackendprotocol with two implementations:MockTrainerBackend— deterministic, network- and GPU-free. ReusesDatasetStore.loadto read/validate the dataset, writes a fake adapter directory (adapter_config.json+ a marker file) and ametrics.jsonwith a fake loss curve. This is the only backend exercised in this project's tests, CLI demo, and CI, on every platform.UnslothTrainerBackend— real Unsloth/TRL/PEFT wiring (lazy imports,FastLanguageModel+ LoRA +SFTTrainer/DPOTrainer). It requires an NVIDIA GPU with CUDA; if imports fail or no CUDA device is found, it raises a clearRuntimeErrorwith install/hardware instructions — it never silently falls back to the mock and never pretends to have trained anything. Honesty note: this repository was built and tested on a Mac with no NVIDIA GPU. Only the import/CUDA guard is tested (tests/test_train_backend.py, via a monkeypatched failing import) — the training loop itself (model loading, PEFT wrapping,SFTTrainer/DPOTrainerexecution, adapter save) has never been run end to end anywhere. Treat it as reviewed-but-unverified code, not a tested path, until someone runs it on real CUDA hardware.
- M3 — Regression eval harness with diff report.
groundedtune eval --before-backend mock|ollama --before-config before.json --after-backend mock|ollama --after-config after.json --eval-set evalsets/core-v1.jsonl [--out report.json](src/groundedtune/eval/). AnEvalBackendprotocol (MockEvalBackendfor canned prompt->response maps,OllamaEvalBackendfor a real local Ollama-compatible/api/chatserver, mirroring assumecheck'sruntime.pypattern — tested only viahttpx.MockTransport, no live server needed). The eval set format is a deliberately simple JSONL of{"prompt", "check": {"type": "contains"|"not_contains"|"regex", "value"}}— no LLM-judge dependency.harness.pyruns a before/after backend pair against the same eval set and buckets results intonewly_passing/newly_failing/stayed_passing/stayed_failing. Unlikeassumecheck'sevalcommand, this one is a CI gate: exit code 0 ifnewly_failingis empty, exit code 1 if there are regressions. - M4 — CLI versioning/rollback.
groundedtune version record|approve|list|show|set-current|rollback(src/groundedtune/versioning.py). A manifest at.groundedtune/versions.jsontracks each recorded version (dataset path + sha256 hash, adapter dir, eval report summary, timestamp,approved: bool, defaultfalse). Critical, tested invariant:set-currentandrollbackrefuse — clean error, non-zero exit — to pointcurrentat a version that isn'tapproved: true(tests/test_versioning.py::test_set_current_on_unapproved_version_is_refusedand the CLI-level equivalent intests/test_cli_version.py).rollbacksetscurrentto the previous approved version before the current one, skipping any unapproved versions in between. This MVP only tracks and gates the pointer — it does not deploy or serve anything.
See examples/m2-m4-walkthrough.sh and
examples/README.md
for a real, captured run of train -> eval -> version record -> approve ->
set-current, including the two steps that are supposed to fail (an
unapproved set-current, a rollback with nothing earlier to roll back
to).
python3 -m venv .venv
.venv/bin/pip install -U pip
.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest -q
.venv/bin/ruff check .MIT — see LICENSE.
