Thinking-Space/One-Shot-OPD

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Python

93

5 commits

updated Sep 4, 2026

See the code

README

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Paper   GitHub   Hugging Face Paper   Hugging Face Collection X Thread

🎉News

  • [2026-09-04] Part II: We examine the role of training data in on-policy distillation (OPD) at the data-minimal limit by training on a single query, and find that OPD is data-overfed but algorithm-starved. Check it out: Paper.
  • [2026-04-15] Part I of this series: Rethinking On-Policy Distillation of Large Language Models.

📖Overview

One-shot OPD versus full-data OPD on mathematical reasoning, and the multi-teacher OPD comparison

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. A single query reaches 71.5% state coverage relative to full-data OPD, with most coverage appearing in the first 100 steps. Adding semantically distinct queries increases state coverage and validation accuracy together, and sixteen queries reach 98.9% coverage and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or all 17k, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.

✨Getting Started

Environment Setup

We implement OPD and MOPD by extending veRL. The vendored copy under verl/ is what the launchers run, so a fresh clone needs no separate veRL install — only its dependencies.

Python 3.12      vllm 0.28.0 (torch 2.13.0  transformers 5.10.4)      verl 0.9.0
git clone https://github.com/Thinking-Space/One-Shot-OPD.git
cd One-Shot-OPD

export MODEL_BASE=/your/models          # student + teacher checkpoints live here
export CONDA_ENV_BIN=/your/env/bin      # vllm >= 0.18, transformers 5.x
  • Reference numbers come from a single 8-GPU node (H100/A100 80 GB); the teacher worker shares the actor's GPU pool.
  • No dataset ships with this repository. Data is rebuilt from public sources by data/prep/; the expected layout, row counts, and data_source values are in data/README.md.

Training

ScriptRewardTeachers
grpo.shrule-based verifier, grouped over n samplesnone
opd.shteacher reverse KLone
mopd.shteacher reverse KL, routed per row by the ability columnseveral
bash recipe/grpo.sh                 # baseline
bash recipe/opd.sh                  # OPD, full data
MODE=oneshot bash recipe/opd.sh     # One-shot OPD
MODE=template bash recipe/opd.sh    # the student writes its own input
bash recipe/mopd.sh                 # MOPD, three teachers

DOMAIN selects the teacher and the parquets, MODE selects which prompts are read, and the two are independent. Anything appended on the command line goes to Hydra.

DOMAIN=code bash recipe/opd.sh              # math (default) / code / fc / if
DOMAIN=fc MODE=oneshot bash recipe/opd.sh   # only math ships a 1-shot corpus
MODE=oneshot LOG_PROB_TOP_K=16 bash recipe/opd.sh trainer.total_training_steps=10

For MOPD, each teacher is a name–path pair passed to multi_teacher_args. Adding one means adding a path variable and extending both the routing map and the call:

export TEACHER_AGENTIC_PATH=${TEACHER_AGENTIC_PATH:-${MODEL_BASE}/MyAgentTeacher}
export ABILITY_TO_TEACHER='{math:math,code:code,instruction_following:if,function_calling:agentic}'

launch "$EXPERIMENT_NAME" \
    $(common_args) \
    $(mode_args) \
    $(multi_teacher_args \
        "math:${TEACHER_MATH_PATH}" \
        "code:${TEACHER_CODE_PATH}" \
        "if:${TEACHER_IF_PATH}" \
        "agentic:${TEACHER_AGENTIC_PATH}") \
    ...
Key parameters
ParameterDefaultDescription
DOMAINmathmath / code / fc / if — selects teacher and datasets
MODEfullfull / oneshot / template — selects the data variant
ACTOR_MODEL_PATHper DOMAINStudent (policy) model to be trained
TEACHER_MODEL_PATHper DOMAINFrozen teacher that provides the token-level reward
MODEL_BASErequiredDirectory the two model paths resolve against
N_RESPONSES1Rollout responses per prompt
LOG_PROB_TOP_K0Top-K tokens kept when computing the token reward; 0 falls back to sampled-token OPD
METRIC_TOP_K0Same computation, observability only — the reward stays plain reverse KL
TOP_K_STRATEGYonly_stuSupport set: only_stu / only_tch / intersection / union / union-intersection
REWARD_WEIGHT_MODEstudent_pToken weighting: student_p / teacher_p / none
VAL_N / VAL_TEMPERATURE / VAL_TOP_P16 / 0.7 / 0.9Validation sampling, giving avg@16

The two top-K knobs mean different things. LOG_PROB_TOP_K changes the training reward; METRIC_TOP_K only measures it.

Default hyperparameters for OPD used in the paper
ItemValue
Rollout batch size64
Mini batch size64
Responses per prompt1
KL coefficient0.0
LogProb top-k0 (default) / 16
Top-k strategystudent top-k
Loss aggregationtoken-mean
Training temperature1.0
Top-p1.0
OptimizerAdamW (beta1 0.9, beta2 0.999, weight decay 0.01)
Learning rate1e-6
Gradient clip norm1.0
Max prompt length1024 (math); 4096 (code, IF, agentic)
Max response length7680 (math, code, IF); 2048 (agentic)

Validation

VAL_ONLY=True bash recipe/opd.sh    # score the current weights, then exit

The in-training validation loop gives avg@16 at temperature 0.7. The evaluation protocols reported in the paper are:

DomainBenchmarksProtocolResponse cap
MathMATH-500, AMC 2023, AIME 2025avg@1631,744
CodeLiveCodeBench v6avg@3, official execution-based evaluator65,536
Instruction followingMulti-IFfinal-turn score averaged over its eight languages16,384 per turn
Agentic tool useBFCL v3avg@8 over the evaluated subsets4,096

Multi-IF

Multi-IF is multi-turn — turn 2's prompt depends on turn 1's answer — which the in-training loop cannot produce. It runs outside verl, and IF training sets VAL_FILES='[]'.

# 8 GPUs, one vLLM engine each (1.5B model, no tensor parallelism)
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
python3 eval/run_if_eval.py multiif --model <hf-dir> --out results/multiif_s50

# Re-score existing generations after a scorer or aggregation change, no GPU
python3 eval/run_if_eval.py rescore --out results/multiif_s50

# Smoke test, a few minutes
python3 eval/run_if_eval.py multiif --model <hf-dir> --out results/smoke --limit 24

📨Contact

🎈Citation

If you find this work helpful, please cite us:

@article{fu2026rethinking,
  title={Rethinking on-policy distillation of large language models ii: One training example},
  author={Fu, Zixuan and He, Bingxiang and Zuo, Yuxin and Huang, Haohuan and Zhang, Jinqian and Xiao, Ruhang and Qian, Cheng and Luo, Qinyu and Gao, Huan-ang and Wang, Yudong and others},
  journal={arXiv preprint arXiv:2609.04172},
  year={2026}
}

License

The code in this repository is released under the Apache License 2.0. The released model checkpoints are subject to the license terms of their respective base models. Please refer to the corresponding model cards for details.

Thinking-Space/One-Shot-OPD

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Python

93

5 commits

updated Sep 4, 2026

See the code

README

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Paper   GitHub   Hugging Face Paper   Hugging Face Collection X Thread

🎉News

  • [2026-09-04] Part II: We examine the role of training data in on-policy distillation (OPD) at the data-minimal limit by training on a single query, and find that OPD is data-overfed but algorithm-starved. Check it out: Paper.
  • [2026-04-15] Part I of this series: Rethinking On-Policy Distillation of Large Language Models.

📖Overview

One-shot OPD versus full-data OPD on mathematical reasoning, and the multi-teacher OPD comparison

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. A single query reaches 71.5% state coverage relative to full-data OPD, with most coverage appearing in the first 100 steps. Adding semantically distinct queries increases state coverage and validation accuracy together, and sixteen queries reach 98.9% coverage and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or all 17k, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.

✨Getting Started

Environment Setup

We implement OPD and MOPD by extending veRL. The vendored copy under verl/ is what the launchers run, so a fresh clone needs no separate veRL install — only its dependencies.

Python 3.12      vllm 0.28.0 (torch 2.13.0  transformers 5.10.4)      verl 0.9.0
git clone https://github.com/Thinking-Space/One-Shot-OPD.git
cd One-Shot-OPD

export MODEL_BASE=/your/models          # student + teacher checkpoints live here
export CONDA_ENV_BIN=/your/env/bin      # vllm >= 0.18, transformers 5.x
  • Reference numbers come from a single 8-GPU node (H100/A100 80 GB); the teacher worker shares the actor's GPU pool.
  • No dataset ships with this repository. Data is rebuilt from public sources by data/prep/; the expected layout, row counts, and data_source values are in data/README.md.

Training

ScriptRewardTeachers
grpo.shrule-based verifier, grouped over n samplesnone
opd.shteacher reverse KLone
mopd.shteacher reverse KL, routed per row by the ability columnseveral
bash recipe/grpo.sh                 # baseline
bash recipe/opd.sh                  # OPD, full data
MODE=oneshot bash recipe/opd.sh     # One-shot OPD
MODE=template bash recipe/opd.sh    # the student writes its own input
bash recipe/mopd.sh                 # MOPD, three teachers

DOMAIN selects the teacher and the parquets, MODE selects which prompts are read, and the two are independent. Anything appended on the command line goes to Hydra.

DOMAIN=code bash recipe/opd.sh              # math (default) / code / fc / if
DOMAIN=fc MODE=oneshot bash recipe/opd.sh   # only math ships a 1-shot corpus
MODE=oneshot LOG_PROB_TOP_K=16 bash recipe/opd.sh trainer.total_training_steps=10

For MOPD, each teacher is a name–path pair passed to multi_teacher_args. Adding one means adding a path variable and extending both the routing map and the call:

export TEACHER_AGENTIC_PATH=${TEACHER_AGENTIC_PATH:-${MODEL_BASE}/MyAgentTeacher}
export ABILITY_TO_TEACHER='{math:math,code:code,instruction_following:if,function_calling:agentic}'

launch "$EXPERIMENT_NAME" \
    $(common_args) \
    $(mode_args) \
    $(multi_teacher_args \
        "math:${TEACHER_MATH_PATH}" \
        "code:${TEACHER_CODE_PATH}" \
        "if:${TEACHER_IF_PATH}" \
        "agentic:${TEACHER_AGENTIC_PATH}") \
    ...
Key parameters
ParameterDefaultDescription
DOMAINmathmath / code / fc / if — selects teacher and datasets
MODEfullfull / oneshot / template — selects the data variant
ACTOR_MODEL_PATHper DOMAINStudent (policy) model to be trained
TEACHER_MODEL_PATHper DOMAINFrozen teacher that provides the token-level reward
MODEL_BASErequiredDirectory the two model paths resolve against
N_RESPONSES1Rollout responses per prompt
LOG_PROB_TOP_K0Top-K tokens kept when computing the token reward; 0 falls back to sampled-token OPD
METRIC_TOP_K0Same computation, observability only — the reward stays plain reverse KL
TOP_K_STRATEGYonly_stuSupport set: only_stu / only_tch / intersection / union / union-intersection
REWARD_WEIGHT_MODEstudent_pToken weighting: student_p / teacher_p / none
VAL_N / VAL_TEMPERATURE / VAL_TOP_P16 / 0.7 / 0.9Validation sampling, giving avg@16

The two top-K knobs mean different things. LOG_PROB_TOP_K changes the training reward; METRIC_TOP_K only measures it.

Default hyperparameters for OPD used in the paper
ItemValue
Rollout batch size64
Mini batch size64
Responses per prompt1
KL coefficient0.0
LogProb top-k0 (default) / 16
Top-k strategystudent top-k
Loss aggregationtoken-mean
Training temperature1.0
Top-p1.0
OptimizerAdamW (beta1 0.9, beta2 0.999, weight decay 0.01)
Learning rate1e-6
Gradient clip norm1.0
Max prompt length1024 (math); 4096 (code, IF, agentic)
Max response length7680 (math, code, IF); 2048 (agentic)

Validation

VAL_ONLY=True bash recipe/opd.sh    # score the current weights, then exit

The in-training validation loop gives avg@16 at temperature 0.7. The evaluation protocols reported in the paper are:

DomainBenchmarksProtocolResponse cap
MathMATH-500, AMC 2023, AIME 2025avg@1631,744
CodeLiveCodeBench v6avg@3, official execution-based evaluator65,536
Instruction followingMulti-IFfinal-turn score averaged over its eight languages16,384 per turn
Agentic tool useBFCL v3avg@8 over the evaluated subsets4,096

Multi-IF

Multi-IF is multi-turn — turn 2's prompt depends on turn 1's answer — which the in-training loop cannot produce. It runs outside verl, and IF training sets VAL_FILES='[]'.

# 8 GPUs, one vLLM engine each (1.5B model, no tensor parallelism)
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
python3 eval/run_if_eval.py multiif --model <hf-dir> --out results/multiif_s50

# Re-score existing generations after a scorer or aggregation change, no GPU
python3 eval/run_if_eval.py rescore --out results/multiif_s50

# Smoke test, a few minutes
python3 eval/run_if_eval.py multiif --model <hf-dir> --out results/smoke --limit 24

📨Contact

🎈Citation

If you find this work helpful, please cite us:

@article{fu2026rethinking,
  title={Rethinking on-policy distillation of large language models ii: One training example},
  author={Fu, Zixuan and He, Bingxiang and Zuo, Yuxin and Huang, Haohuan and Zhang, Jinqian and Xiao, Ruhang and Qian, Cheng and Luo, Qinyu and Gao, Huan-ang and Wang, Yudong and others},
  journal={arXiv preprint arXiv:2609.04172},
  year={2026}
}

License

The code in this repository is released under the Apache License 2.0. The released model checkpoints are subject to the license terms of their respective base models. Please refer to the corresponding model cards for details.