Skip to content
#

grpo-training

Here are 14 public repositories matching this topic...

Train a small reasoning model to obey an explicit thinking-token budget ("Think for maximum N tokens") using GRPO. Reproduces the L1/LCPO recipe (Aggarwal & Welleck, CMU 2025) on DeepSeek-R1-Distill-Qwen-1.5B, then extends it with a GGUF release and a llama.cpp demo of the budget knob.

  • Updated Sep 8, 2026
  • Python

Add this topic to your repo

To associate your repository with the grpo-training topic, visit your repo's landing page and select "manage topics."

Learn more