Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper
Python
20
2 commits
updated Jul 4, 2025
This repository contains the codebase and resources for the paper:
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
This work argues for evaluating large language models (LLMs) using free-form answer matching instead of traditional multiple-choice (MCQ) formats. We provide code, datasets, and tools for reproducing our experiments, training classifiers, and analyzing results across a variety of benchmarks.
This repo uses uv as a fast, modern Python package manager. To get started:
# Install uv if you don't have it
pip install uv
# Create and activate a virtual environment
uv venv qa
source qa/bin/activate
# Install PyTorch (CUDA 12.1)
uv pip install torch --index-url https://download.pytorch.org/whl/cu121
# Install all dependencies for lmeval (API, wandb, vllm, dev extras)
cd lmeval
uv pip install -e ".[api,wandb,vllm,dev]"
mcq_classifier/: Code to train a classifier on choices-only data from MCQ benchmarks (e.g., MMLU Pro, SuperGQPA, YourBench, TruthfulQA, HellaSwag, ARC) using HuggingFace Transformers and Accelerate. Includes both language and vision MCQ scripts. See its README for details.
src/query_models/: Scripts for querying various language models with different question types (free-form, MCQ) and extracting their responses. Supports datasets like GPQA and MMLU-Pro.
src/filtering/: Tools for filtering and combining QA datasets, including scripts for cleaning, merging, and scoring questions and answers.
src/cost_analysis/: Tools for analyzing computational costs and token usage across different evaluation methods, including cost breakdowns and visualizations.
src/judge_w_gt/: Framework for evaluating model responses using LLMs as matchers or judges, supporting both ground-truth (with reference answer) and free-form (without reference answer) judging for various question types.
src/visualize_resps/: Flask-based annotation interface for labeling model responses; used for human annotation and navigation of QA data.
We provide all free-form datasets, human annotations, and sample-level model outputs evaluated in our paper as a HuggingFace collection:
This collection includes:
If you use these resources, please cite our paper. For questions or contributions, open an issue or pull request!
Python
68.4%
Jupyter Notebook
29.6%
HTML
1.6%
Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper
Python
20
2 commits
updated Jul 4, 2025
This repository contains the codebase and resources for the paper:
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
This work argues for evaluating large language models (LLMs) using free-form answer matching instead of traditional multiple-choice (MCQ) formats. We provide code, datasets, and tools for reproducing our experiments, training classifiers, and analyzing results across a variety of benchmarks.
This repo uses uv as a fast, modern Python package manager. To get started:
# Install uv if you don't have it
pip install uv
# Create and activate a virtual environment
uv venv qa
source qa/bin/activate
# Install PyTorch (CUDA 12.1)
uv pip install torch --index-url https://download.pytorch.org/whl/cu121
# Install all dependencies for lmeval (API, wandb, vllm, dev extras)
cd lmeval
uv pip install -e ".[api,wandb,vllm,dev]"
mcq_classifier/: Code to train a classifier on choices-only data from MCQ benchmarks (e.g., MMLU Pro, SuperGQPA, YourBench, TruthfulQA, HellaSwag, ARC) using HuggingFace Transformers and Accelerate. Includes both language and vision MCQ scripts. See its README for details.
src/query_models/: Scripts for querying various language models with different question types (free-form, MCQ) and extracting their responses. Supports datasets like GPQA and MMLU-Pro.
src/filtering/: Tools for filtering and combining QA datasets, including scripts for cleaning, merging, and scoring questions and answers.
src/cost_analysis/: Tools for analyzing computational costs and token usage across different evaluation methods, including cost breakdowns and visualizations.
src/judge_w_gt/: Framework for evaluating model responses using LLMs as matchers or judges, supporting both ground-truth (with reference answer) and free-form (without reference answer) judging for various question types.
src/visualize_resps/: Flask-based annotation interface for labeling model responses; used for human annotation and navigation of QA data.
We provide all free-form datasets, human annotations, and sample-level model outputs evaluated in our paper as a HuggingFace collection:
This collection includes:
If you use these resources, please cite our paper. For questions or contributions, open an issue or pull request!
Python
68.4%
Jupyter Notebook
29.6%
HTML
1.6%