Yuqian-Fu/SRFT-Qwen2.5-Math-7B

Model

πŸ“„ Introduction

3

9 commits

1 linked in READMEs

updated Jul 24, 2025

See the code

README

πŸ“„ Introduction

Supervised Reinforcement Fine-Tuning (SRFT) is a single-stage method that unifies both fine-tuning paradigms through entropy-aware weighting mechanisms.

Paper: arXiv

Project Website: SRFT

conversational
endpoints_compatible
qwen2
safetensors
text-generation
text-generation-inference
transformers

Contributors

Yuqian-Fu

5 commits

nielsr

2 commits

YF
Yuqian Fu

2 commits

Yuqian-Fu/SRFT-Qwen2.5-Math-7B

Model

πŸ“„ Introduction

3

9 commits

1 linked in READMEs

updated Jul 24, 2025

See the code

README

πŸ“„ Introduction

Supervised Reinforcement Fine-Tuning (SRFT) is a single-stage method that unifies both fine-tuning paradigms through entropy-aware weighting mechanisms.

Paper: arXiv

Project Website: SRFT

conversational
endpoints_compatible
qwen2
safetensors
text-generation
text-generation-inference
transformers

Contributors

Yuqian-Fu

5 commits

nielsr

2 commits

YF
Yuqian Fu

2 commits