benjaminzwhite/reasoning-models

Experiments with reasoning models, training techniques, papers

30

731 commits

updated Sep 22, 2026

See the code

README

Reasoning

Reading list and comments/short reviews around reasoning models, training techniques, and related papers

TODO

  • build tag system e.g. Policies {DPO, GRPO, ...} with a view to building reading list, integrate with Zotero/Obsidian
  • use new approach for formatting table and generating summary: test time and cost first

Papers

Priority

Topic - General or unsorted

Topic - Audio Reasoning

Topic - Routing

Topic - Reasoning in Diffusion Models

Topic - Parallel Reasoning

Start with the first-linked survey.

Topic - Verifier-free RL and approaches without External Rewards

Emerging topic, seen a few papers around this theme recently.

Topic - Surveys and reviews

Topic - Tool Use/Function Calling

Topic - Reasoning + Agents

Topic - X of Thoughts variants

Topic - Efficiency, Decoding Strategies, Implementation Tricks etc.

Topic - Ensembling, Boosting, Stacking etc.

Topic - Evaluation of reasoning

Topic - Mathematics

Topic - Coding

Topic - Biology

Topic - Multimodal

TODO: expand with separate tags

Topic - Robotics

Personal reading notes/papers that gave me some reasoning-related ideas


Datasets


Benchmarks


Blogs and articles

Workshops


Repos

Collections

benchmarks
cot
deepseek
deepseek-r1
llm
models
o1
papers
reasoning
reinforcement-learning
research
rl

Contributors

benjaminzwhite

731 commits

benjaminzwhite/reasoning-models

Experiments with reasoning models, training techniques, papers

30

731 commits

updated Sep 22, 2026

See the code

README

Reasoning

Reading list and comments/short reviews around reasoning models, training techniques, and related papers

TODO

  • build tag system e.g. Policies {DPO, GRPO, ...} with a view to building reading list, integrate with Zotero/Obsidian
  • use new approach for formatting table and generating summary: test time and cost first

Papers

Priority

Topic - General or unsorted

Topic - Audio Reasoning

Topic - Routing

Topic - Reasoning in Diffusion Models

Topic - Parallel Reasoning

Start with the first-linked survey.

Topic - Verifier-free RL and approaches without External Rewards

Emerging topic, seen a few papers around this theme recently.

Topic - Surveys and reviews

Topic - Tool Use/Function Calling

Topic - Reasoning + Agents

Topic - X of Thoughts variants

Topic - Efficiency, Decoding Strategies, Implementation Tricks etc.

Topic - Ensembling, Boosting, Stacking etc.

Topic - Evaluation of reasoning

Topic - Mathematics

Topic - Coding

Topic - Biology

Topic - Multimodal

TODO: expand with separate tags

Topic - Robotics

Personal reading notes/papers that gave me some reasoning-related ideas


Datasets


Benchmarks


Blogs and articles

Workshops


Repos

Collections

benchmarks
cot
deepseek
deepseek-r1
llm
models
o1
papers
reasoning
reinforcement-learning
research
rl

Contributors

benjaminzwhite

731 commits