LLM Decoding & Math Reasoning

11 repos

Techniques for improving language model output quality through advanced decoding strategies and mathematical reasoning verification. The cluster centers on beam search, best-of-N sampling, and reward model-based selection methods applied to instruction-tuned Llama models, with particular emphasis on mathematical problem-solving and answer verification. Repositories include both the decoding implementations themselves and curated datasets of model completions with quality scores for training and evaluation.