Synthetic-document fine-tuning on Qwen2.5-7B: a controlled study of whether SDF installs sandbagging, finding a layered recognition/generation/behavior dissociation.
-
Updated
May 25, 2026 - Python
Synthetic-document fine-tuning on Qwen2.5-7B: a controlled study of whether SDF installs sandbagging, finding a layered recognition/generation/behavior dissociation.
Reproducible red-team findings for openai/gpt-oss-20b: five minimal harnesses with checks, zips & manifest (v0.9.3).
Proving that the neural network is honest about its lack of capabilities
Preregistered AI-safety study of sandbagging model organisms: cue-sharing, not trigger syntax, determines whether two locked capabilities share an unlock direction. Password locking destroys pre-existing cross-capability alignment; semantic locking amplifies it. All five predictions failed, four reversed.
To associate your repository with the sandbagging topic, visit your repo's landing page and select "manage topics."