[ICML 2026] Reasoning in Parallelism via Self-Distilled RL
Python
113
45 commits
updated Jun 28, 2026
We introduce the Native Parallel Reasoner (NPR), a scalable framework for constructing models that intrinsically reason in parallelism. NPR learns adaptive decomposition and aggregation policies through a teacher-free pipeline combining self-distilled parallel Supervised Fine-Tuning (SFT) with Native Parallel Reinforcement Learning (RL). This approach allows the model to optimize its own branching strategies directly from experience within a shared computation graph, preserving its native reasoning style while maximizing exploration efficiency. Across eight diverse reasoning benchmarks, NPR achieves decisive gains: self-distilled data outperform prior teacher-generated corpora by 10.1%, and our Parallel RL stage improves over direct RL baselines by 3.0%. Crucially, NPR delivers up to 4.6× inference acceleration over autoregressive baselines and exhibits genuine, non-simulated parallel reasoning behaviors.
# Create env for NPR-Zero
cd npr-zero
conda create -n zero python=3.11
conda activate zero
conda install nvidia::cuda-nvcc
# Install dependencies
pip install -e .[sglang]
pip install liger-kernel
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
pip uninstall pynvml
pip install "latex2sympy2-extended[antlr4_9_3]"
experiments/raw_data folder.python examples/data_preprocess/orz.pypython examples/data_preprocess/aime25.pyModify the RAY_DATA_HOME and MODEL_PATH to yours.
# Run NPR-Zero training
bash experiments/run.sh
# Create env for NPR-Beta
cd npr-Beta
conda create -n warmup python=3.11 -y
conda activate warmup
# Install dependencies
pip install -r requirements.txt
# Perform rejection sampling
bash scripts/sampling.sh
Key parameters in sampling.sh:
MODEL_PATH: Path to model checkpoint (Stage 1)OUTPUT_DIR: Output directory for sampled trajectories--dataset: Dataset name (default: ORZ-MATH-57K)--instruction: Prompt template file--max_sample_trial: Max sampling attempts per problem (default: 8)--temperature: Sampling temperature (default: 1.0)# Start warmup training
bash train/sft_math.sh
Key parameters in sft_math.sh:
base_model: Base model to fine-tune (default: Qwen3-4B-Instruct)train_file_path: Training trajectories path (default: dataset/math/rejection_sampling/train)lr: Learning rateepochs: Number of training epochsckpts/NPR-Warmup-4B-Inst-{timestamp}/# Create env for NPR-RL
cd npr-rl
conda create -n rl python=3.11
conda activate rl
conda install -c nvidia cuda-nvcc=12.8
# Install dependencies
cd npr-rl
pip install -e .
pip install liger-kernel
pip install "latex2sympy2-extended[antlr4_9_3]"
cd verl/workers/rollout/sglang_rollout/sglang/python
pip install -e .[all]
pip install transformers==4.53.1
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
pip install fire
pip uninstall pynvml
experiments/raw_data folder.python examples/data_preprocess/orz.pypython examples/data_preprocess/aime.pyModify the RAY_DATA_HOME and MODEL_PATH to yours.
Note the MODEL_PATH is from Stage 2.
# Run native parallel RL
bash experiments/run.sh
For AR:
# Create env for evaluations
cd evals
conda create -n eval python=3.10
conda activate eval
# Install dependencies
pip install -r requirements.txt
For NPR:
# Inheriting the environment of the third stage.
conda activate rl
pip install word2number rich latex2sympy2==1.9.1
python convert_to_hf.py verl/experiments/ckpts/project_name/exp_name/global_step_x/actor <STAGE_2_MODEL_PATH> <TARGET_HF_MODEL_PATH>Modify the <<TARGET_HF_MODEL_PATH>> to yours.
# Start evaluation of AIME25
./scripts/eval.sh \
--cuda 0,1,2,3,4,5,6,7 \
--tp_size 2 \
--dp_size 4 \
--task "AIME25" \
--max_eval_samples 30 \
--eval_batch_size 8 \
--model_path <TARGET_HF_MODEL_PATH> \
--prompt_path prompts/npr.txt \
--engine parallel \
--num_samples 1 \
--k 1 \
--max_new_tokens 40000 \
--temperature 1.0 \
--top_p 0.7 \
--top_k -1 \
--overwrite \
--apply_chat
We report Pass@1 accuracy averaged over 8 samples for each problem as below.
Native-Parallel-Reasoner project.git clone https://github.com/bigai-nlco/Native-Parallel-Reasoner.git
git checkout -b your_name/feature-x
CE```
git commit -m 'Implemented new feature x.'
git push origin feature-x
Native Parallel Reasoner is protected under the LICENSE. For more details, please refer to the LICENSE file.
@inproceedings{
wu2026native,
title={Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning},
author={Tong Wu and Yang Liu and Jun Bai and Zixia Jia and Shuyi Zhang and Ziyong Lin and Yanting Wang and Song-Chun Zhu and Zilong Zheng},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=tyqW6SYxWB}
}
This codebase is influenced by these remarkable projects of AI community, including verl, sglang, and Multiverse.
Python
86.5%
Shell
5.5%
C++
2.2%
Cuda
2.1%
Jupyter Notebook
1.7%
[ICML 2026] Reasoning in Parallelism via Self-Distilled RL
Python
113
45 commits
updated Jun 28, 2026
We introduce the Native Parallel Reasoner (NPR), a scalable framework for constructing models that intrinsically reason in parallelism. NPR learns adaptive decomposition and aggregation policies through a teacher-free pipeline combining self-distilled parallel Supervised Fine-Tuning (SFT) with Native Parallel Reinforcement Learning (RL). This approach allows the model to optimize its own branching strategies directly from experience within a shared computation graph, preserving its native reasoning style while maximizing exploration efficiency. Across eight diverse reasoning benchmarks, NPR achieves decisive gains: self-distilled data outperform prior teacher-generated corpora by 10.1%, and our Parallel RL stage improves over direct RL baselines by 3.0%. Crucially, NPR delivers up to 4.6× inference acceleration over autoregressive baselines and exhibits genuine, non-simulated parallel reasoning behaviors.
# Create env for NPR-Zero
cd npr-zero
conda create -n zero python=3.11
conda activate zero
conda install nvidia::cuda-nvcc
# Install dependencies
pip install -e .[sglang]
pip install liger-kernel
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
pip uninstall pynvml
pip install "latex2sympy2-extended[antlr4_9_3]"
experiments/raw_data folder.python examples/data_preprocess/orz.pypython examples/data_preprocess/aime25.pyModify the RAY_DATA_HOME and MODEL_PATH to yours.
# Run NPR-Zero training
bash experiments/run.sh
# Create env for NPR-Beta
cd npr-Beta
conda create -n warmup python=3.11 -y
conda activate warmup
# Install dependencies
pip install -r requirements.txt
# Perform rejection sampling
bash scripts/sampling.sh
Key parameters in sampling.sh:
MODEL_PATH: Path to model checkpoint (Stage 1)OUTPUT_DIR: Output directory for sampled trajectories--dataset: Dataset name (default: ORZ-MATH-57K)--instruction: Prompt template file--max_sample_trial: Max sampling attempts per problem (default: 8)--temperature: Sampling temperature (default: 1.0)# Start warmup training
bash train/sft_math.sh
Key parameters in sft_math.sh:
base_model: Base model to fine-tune (default: Qwen3-4B-Instruct)train_file_path: Training trajectories path (default: dataset/math/rejection_sampling/train)lr: Learning rateepochs: Number of training epochsckpts/NPR-Warmup-4B-Inst-{timestamp}/# Create env for NPR-RL
cd npr-rl
conda create -n rl python=3.11
conda activate rl
conda install -c nvidia cuda-nvcc=12.8
# Install dependencies
cd npr-rl
pip install -e .
pip install liger-kernel
pip install "latex2sympy2-extended[antlr4_9_3]"
cd verl/workers/rollout/sglang_rollout/sglang/python
pip install -e .[all]
pip install transformers==4.53.1
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
pip install fire
pip uninstall pynvml
experiments/raw_data folder.python examples/data_preprocess/orz.pypython examples/data_preprocess/aime.pyModify the RAY_DATA_HOME and MODEL_PATH to yours.
Note the MODEL_PATH is from Stage 2.
# Run native parallel RL
bash experiments/run.sh
For AR:
# Create env for evaluations
cd evals
conda create -n eval python=3.10
conda activate eval
# Install dependencies
pip install -r requirements.txt
For NPR:
# Inheriting the environment of the third stage.
conda activate rl
pip install word2number rich latex2sympy2==1.9.1
python convert_to_hf.py verl/experiments/ckpts/project_name/exp_name/global_step_x/actor <STAGE_2_MODEL_PATH> <TARGET_HF_MODEL_PATH>Modify the <<TARGET_HF_MODEL_PATH>> to yours.
# Start evaluation of AIME25
./scripts/eval.sh \
--cuda 0,1,2,3,4,5,6,7 \
--tp_size 2 \
--dp_size 4 \
--task "AIME25" \
--max_eval_samples 30 \
--eval_batch_size 8 \
--model_path <TARGET_HF_MODEL_PATH> \
--prompt_path prompts/npr.txt \
--engine parallel \
--num_samples 1 \
--k 1 \
--max_new_tokens 40000 \
--temperature 1.0 \
--top_p 0.7 \
--top_k -1 \
--overwrite \
--apply_chat
We report Pass@1 accuracy averaged over 8 samples for each problem as below.
Native-Parallel-Reasoner project.git clone https://github.com/bigai-nlco/Native-Parallel-Reasoner.git
git checkout -b your_name/feature-x
CE```
git commit -m 'Implemented new feature x.'
git push origin feature-x
Native Parallel Reasoner is protected under the LICENSE. For more details, please refer to the LICENSE file.
@inproceedings{
wu2026native,
title={Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning},
author={Tong Wu and Yang Liu and Jun Bai and Zixia Jia and Shuyi Zhang and Ziyong Lin and Yanting Wang and Song-Chun Zhu and Zilong Zheng},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=tyqW6SYxWB}
}
This codebase is influenced by these remarkable projects of AI community, including verl, sglang, and Multiverse.
Python
86.5%
Shell
5.5%
C++
2.2%
Cuda
2.1%
Jupyter Notebook
1.7%