
The AmirhoseinGH/DS-Qwen-1.5b-GG-CalibratedConfRL model is derived from the DeepSeek-R1 Qwen Distill 1.5B base model, optimized through confidence-based Reinforcement Learning using GRPO for enhanced intrinsic confidence calibration.
This calibration process significantly improves the reliability of the model’s internal confidence signals. The model is optimized for use with the Guided by Gut (GG) framework, a self-guided test-time scaling (TTS) strategy that leverages these intrinsic confidence signals to perform complex reasoning tasks efficiently—without costly external verifier models.
Traditional TTS methods often require substantial computational resources due to their reliance on external verifier models like Process Reward Models (PRMs) or extensive sampling strategies (e.g., Best-of-N). The GG framework provides a powerful yet computationally efficient alternative:
The RL fine-tuning is computationally efficient and minimal:
5 commits
1 commits

The AmirhoseinGH/DS-Qwen-1.5b-GG-CalibratedConfRL model is derived from the DeepSeek-R1 Qwen Distill 1.5B base model, optimized through confidence-based Reinforcement Learning using GRPO for enhanced intrinsic confidence calibration.
This calibration process significantly improves the reliability of the model’s internal confidence signals. The model is optimized for use with the Guided by Gut (GG) framework, a self-guided test-time scaling (TTS) strategy that leverages these intrinsic confidence signals to perform complex reasoning tasks efficiently—without costly external verifier models.
Traditional TTS methods often require substantial computational resources due to their reliance on external verifier models like Process Reward Models (PRMs) or extensive sampling strategies (e.g., Best-of-N). The GG framework provides a powerful yet computationally efficient alternative:
The RL fine-tuning is computationally efficient and minimal:
5 commits
1 commits