A transparent, single-file implementation for understanding GRPO (K1 in Rewards), free from the abstractions of large libraries.
Python
91
6 commits
updated Feb 14, 2026
A minimal, single-file implementation of Group Relative Policy Optimization (GRPO).
Current RLHF/RL libraries (TRL, OpenRLHF) are powerful but often hide the mathematical logic behind layers of abstraction, callbacks, and complex class inheritances. Transparent GRPO is designed for researchers and engineers who want to:
loss is actually calculated.Generate -> Reward -> Advantage -> Update.torch, transformers, accelerate, and numpy. No heavy RL frameworks required.transparent_grpo.py: The main script containing the entire GRPO implementation. Read from top to bottom to understand the flow.logs/train_grpo_20260209_172905.log: my example training log showing the reward and loss progression.No other things to worry about. No hidden files, no complex directory structure. Just one file to read and understand.
Estimated VRAM usage for Qwen2.5-3B-Instruct (used in the code) with group_size=4, max_new_tokens=512, and DeepSpeed ZeRO-2:
| GPU Setup | Estimated VRAM per GPU | Status | Note |
|---|---|---|---|
| 1x A100 (80GB) | ~35 GB | β Easy | Perfect for single-GPU debugging. |
| 2x RTX 3090/4090 (24GB) | ~22 GB | β οΈ Tight | Requires ZeRO-2 to shard optimizer states. |
| 4x L4 (24GB) | ~16 GB | β Safe | Optimizer states are well-distributed. |
| 1x RTX 4090 (24GB) | OOM | β Fail | OOM unless you enable ZeRO-Offload (CPU). |
You only need standard PyTorch ecosystem libraries:
pip install torch transformers accelerate deepspeed numpy
The script is configured to work out-of-the-box with Hugging Face Accelerate and DeepSpeed (ZeRO-2 recommended for the 3B model).
# Configure accelerate first if you haven't (select DeepSpeed/Zero2)
accelerate config
# Run the training script
nohup stdbuf -oL accelerate launch --use_deepspeed --zero_stage 2 transparent_grpo.py > logs/train_grpo_$(date +%Y%m%d_%H%M%S).log 2>&1 &
This repository strips away the magic. Here is where the core GRPO mechanics live in transparent_grpo.py:
| Component | Description |
|---|---|
| Generation | Standard model.generate() is called inside the loop. No hidden Actor classes. |
| Group Sampling | Config.group_size = 4. We generate $G$ outputs for every prompt. |
| Reward Function | A custom regular-expression based reward system for a specific algebra problem (ToyEnv in code). |
| Ref Model LogProbs | Calculate $\pi_{\text{ref}}(y | x)$ |
| KL Divergence | Calculate the per-token KL: $\log \pi_\theta(y | x) - \log \pi_{\text{ref}}(y | x)$ |
| The Advantage | Calculated simply as (rewards - mean) / std. |
| Loss Function | Standard PPO clipping: $\min(r_t A_t, \text{clip}(r_t, 1-\epsilon, 1+\epsilon) A_t)$. |
Compared with DeepSeekMath (Original GRPO paper), this implementation:
Unlike older PPO implementations that put KL divergence in the loss function, this implementation follows Shah et al., 2026 that subtracts the KL penalty directly from the rewards:
# Code snippet from transparent_grpo.py
per_token_kl = token_log_probs.detach() - ref_token_log_probs.detach()
kl_penalty = (per_token_kl * loss_mask).sum(dim=1)
rewards_with_kl = rewards - Config.beta * kl_penalty
This approach (Unbiased K1 Estimator) provides better var & bias balanced gradients (more stable training) and better performance.

Image from Shah et al., 2026 showing the variance reduction benefits of K1 in Rewards vs K3 in Loss.
ToyEnvTo ensure the script runs quickly on modest hardware, it solves one specific Piecewise Function Continuity problem.
Note: This acts as a unit test for the algorithm. If the loss drops and reward goes to 1.0, the GRPO implementation is correct.
This implementation prioritizes readability and educational value over production efficiency.
generate(). For high-throughput training, integrate vLLM or SGLang.Python
100.0%
A transparent, single-file implementation for understanding GRPO (K1 in Rewards), free from the abstractions of large libraries.
Python
91
6 commits
updated Feb 14, 2026
A minimal, single-file implementation of Group Relative Policy Optimization (GRPO).
Current RLHF/RL libraries (TRL, OpenRLHF) are powerful but often hide the mathematical logic behind layers of abstraction, callbacks, and complex class inheritances. Transparent GRPO is designed for researchers and engineers who want to:
loss is actually calculated.Generate -> Reward -> Advantage -> Update.torch, transformers, accelerate, and numpy. No heavy RL frameworks required.transparent_grpo.py: The main script containing the entire GRPO implementation. Read from top to bottom to understand the flow.logs/train_grpo_20260209_172905.log: my example training log showing the reward and loss progression.No other things to worry about. No hidden files, no complex directory structure. Just one file to read and understand.
Estimated VRAM usage for Qwen2.5-3B-Instruct (used in the code) with group_size=4, max_new_tokens=512, and DeepSpeed ZeRO-2:
| GPU Setup | Estimated VRAM per GPU | Status | Note |
|---|---|---|---|
| 1x A100 (80GB) | ~35 GB | β Easy | Perfect for single-GPU debugging. |
| 2x RTX 3090/4090 (24GB) | ~22 GB | β οΈ Tight | Requires ZeRO-2 to shard optimizer states. |
| 4x L4 (24GB) | ~16 GB | β Safe | Optimizer states are well-distributed. |
| 1x RTX 4090 (24GB) | OOM | β Fail | OOM unless you enable ZeRO-Offload (CPU). |
You only need standard PyTorch ecosystem libraries:
pip install torch transformers accelerate deepspeed numpy
The script is configured to work out-of-the-box with Hugging Face Accelerate and DeepSpeed (ZeRO-2 recommended for the 3B model).
# Configure accelerate first if you haven't (select DeepSpeed/Zero2)
accelerate config
# Run the training script
nohup stdbuf -oL accelerate launch --use_deepspeed --zero_stage 2 transparent_grpo.py > logs/train_grpo_$(date +%Y%m%d_%H%M%S).log 2>&1 &
This repository strips away the magic. Here is where the core GRPO mechanics live in transparent_grpo.py:
| Component | Description |
|---|---|
| Generation | Standard model.generate() is called inside the loop. No hidden Actor classes. |
| Group Sampling | Config.group_size = 4. We generate $G$ outputs for every prompt. |
| Reward Function | A custom regular-expression based reward system for a specific algebra problem (ToyEnv in code). |
| Ref Model LogProbs | Calculate $\pi_{\text{ref}}(y | x)$ |
| KL Divergence | Calculate the per-token KL: $\log \pi_\theta(y | x) - \log \pi_{\text{ref}}(y | x)$ |
| The Advantage | Calculated simply as (rewards - mean) / std. |
| Loss Function | Standard PPO clipping: $\min(r_t A_t, \text{clip}(r_t, 1-\epsilon, 1+\epsilon) A_t)$. |
Compared with DeepSeekMath (Original GRPO paper), this implementation:
Unlike older PPO implementations that put KL divergence in the loss function, this implementation follows Shah et al., 2026 that subtracts the KL penalty directly from the rewards:
# Code snippet from transparent_grpo.py
per_token_kl = token_log_probs.detach() - ref_token_log_probs.detach()
kl_penalty = (per_token_kl * loss_mask).sum(dim=1)
rewards_with_kl = rewards - Config.beta * kl_penalty
This approach (Unbiased K1 Estimator) provides better var & bias balanced gradients (more stable training) and better performance.

Image from Shah et al., 2026 showing the variance reduction benefits of K1 in Rewards vs K3 in Loss.
ToyEnvTo ensure the script runs quickly on modest hardware, it solves one specific Piecewise Function Continuity problem.
Note: This acts as a unit test for the algorithm. If the loss drops and reward goes to 1.0, the GRPO implementation is correct.
This implementation prioritizes readability and educational value over production efficiency.
generate(). For high-throughput training, integrate vLLM or SGLang.Python
100.0%