Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.
Python
42
1,103 commits
updated Oct 3, 2026
Comprehensive LLM evaluation at scale - A production-ready framework for evaluating large language models across 90+ benchmarks.

For full documentation, visit our Docs Page.
The codebase is tested and compatible with Python 3.12 and PyTorch 2.5. You will also need the appropriate CUDA dependencies and version installed on your system for GPU support. Detailed installation instructions can be found here.
The easiest way to get started is by installing the library via pip and use it as an external dependency.
pip install eval_framework
There are optional extras available to unlock specific features of the library:
api for inference using the aleph-alpha client.determined for running jobs via determined.openai for inference against OpenAI-compatible HTTP endpoints.transformers for inference using the transformers library.As a short hand, the all extra installs all of the above.
We use uv to better resolve dependencies when downloading the extras. You can install uv with:
curl -LsSf https://astral.sh/uv/install.sh | sh
or by follwing the uv installation docs.
Now, you can safely install the project with all optional extras:
uv sync --all-extras
or with pip
uv pip install eval_framework[all]
Tip: ensure python is properly installed with uv:
uv python install 3.12 --reinstall
We provide custom groups to control optional extras.
flash_attn: Install flash_attn with correct handling of build isolationThus, the following will setup the project with flash_attn
uv sync --all-extras --group flash_attn
To evaluate a single benchmark locally, you can use the following command:
eval_framework \
--models src/eval_framework/llm/models.py \
--llm-name Smollm135MInstruct \
--task-name "MMLU" \
--task-subjects "abstract_algebra" \
--output-dir ./eval_results \
--num-fewshot 5 \
--num-samples 10
For more detailed CLI usage instructions, see the CLI Usage Guide.
Subset of core capabilities benchmarks coverd by eval-framework:
| Reasoning | Knowledge | Math | Coding | Structured outputs | Long Context |
|---|---|---|---|---|---|
| COPA, BalancedCOPA | ARC | AIME | BigCodeBench | IFEval | InfiniteBench |
| Hellaswag | MMLU | GSM8K | HumanEval | StructEval | QUALITY |
| Winogrande | Openbook QA | MATH-500 | MBPP | ZeroSCROLLS |
Subset of language-specific and domain-specific benchmarks coverd by eval-framework:
| Multilingual | Specialized | Safety & Bias | Efficiency Metrics |
|---|---|---|---|
| WMT Translation | MMLU | TruthfulQA | Compression ratios |
| FLORES-200 | Legal (CaseHold) | Winogender | Runtime |
| Multilingual MMLU | Scientific (SciQ) | ||
| German/Finnish tasks |
Tasks focused on logical reasoning, text distillation, instruction following, and output control. Examples include:
Tasks emphasizing classification, reasoning, and open QA. Examples include:
Tasks designed for long-context scenarios, including QA, summarization, and aggregation. Examples include:
Evaluation metrics include:
For the full list of tasks and metrics, see Detailed Task Table.
Eval-Framework provides a unified interface for evaluating language models across diverse benchmarks. The framework follows this interaction model:
BaseLLM interface (HuggingFace, OpenAI, custom APIs)BaseTask (completion, loglikelihood, or LLM-judge based)BaseMetric classespip install eval_framework[transformers]
from functools import partial
from pathlib import Path
from eval_framework.llm.huggingface import HFLLM
from eval_framework.main import main
from eval_framework.tasks.eval_config import EvalConfig
from template_formatting.formatter import HFFormatter
# Define your model
class MyHuggingFaceModel(HFLLM):
LLM_NAME = "microsoft/DialoGPT-medium"
DEFAULT_FORMATTER = partial(HFFormatter, "microsoft/DialoGPT-medium")
if __name__ == "__main__":
# Initialize your model
llm = MyHuggingFaceModel()
# Running evaluation on MMLU abstract algebra task using 5 few-shot examples and 10 samples
config = EvalConfig(
output_dir=Path("./eval_results"),
num_fewshot=5,
num_samples=10,
task_name="MMLU",
task_subjects=["abstract_algebra", "astronomy"],
llm_class=MyHuggingFaceModel,
)
# Run evaluation and get results
results = main(llm=llm, config=config)
./eval_results/ for detailed outputs and use our results guide to interpret themIf you use eval-framework in your research, please cite:
@software{eval_framework,
author={Aleph Alpha Research},
title={Aleph Alpha Eval Framework},
year={2026},
version = {x.y.z},
url={https://github.com/Aleph-Alpha-Research/eval-framework}
}
This project is licensed under the Apache License 2.0.
The constituent tasks' datasets can be subject to specific, more restrictive license terms by third parties. By using eval-framework you agree to adhere to and be bound by such license terms in the individual case.
This project has received funding from the European Union’s Digital Europe Programme under grant agreement No. 101195233 (OpenEuroLLM).
The contents of this publication are the sole responsibility of the OpenEuroLLM consortium and do not necessarily reflect the opinion of the European Union.
Python
100.0%
Comprehensive LLM evaluation at scale: A production-ready framework for evaluating large language models across multiple benchmarks.
Python
42
1,103 commits
updated Oct 3, 2026
Comprehensive LLM evaluation at scale - A production-ready framework for evaluating large language models across 90+ benchmarks.

For full documentation, visit our Docs Page.
The codebase is tested and compatible with Python 3.12 and PyTorch 2.5. You will also need the appropriate CUDA dependencies and version installed on your system for GPU support. Detailed installation instructions can be found here.
The easiest way to get started is by installing the library via pip and use it as an external dependency.
pip install eval_framework
There are optional extras available to unlock specific features of the library:
api for inference using the aleph-alpha client.determined for running jobs via determined.openai for inference against OpenAI-compatible HTTP endpoints.transformers for inference using the transformers library.As a short hand, the all extra installs all of the above.
We use uv to better resolve dependencies when downloading the extras. You can install uv with:
curl -LsSf https://astral.sh/uv/install.sh | sh
or by follwing the uv installation docs.
Now, you can safely install the project with all optional extras:
uv sync --all-extras
or with pip
uv pip install eval_framework[all]
Tip: ensure python is properly installed with uv:
uv python install 3.12 --reinstall
We provide custom groups to control optional extras.
flash_attn: Install flash_attn with correct handling of build isolationThus, the following will setup the project with flash_attn
uv sync --all-extras --group flash_attn
To evaluate a single benchmark locally, you can use the following command:
eval_framework \
--models src/eval_framework/llm/models.py \
--llm-name Smollm135MInstruct \
--task-name "MMLU" \
--task-subjects "abstract_algebra" \
--output-dir ./eval_results \
--num-fewshot 5 \
--num-samples 10
For more detailed CLI usage instructions, see the CLI Usage Guide.
Subset of core capabilities benchmarks coverd by eval-framework:
| Reasoning | Knowledge | Math | Coding | Structured outputs | Long Context |
|---|---|---|---|---|---|
| COPA, BalancedCOPA | ARC | AIME | BigCodeBench | IFEval | InfiniteBench |
| Hellaswag | MMLU | GSM8K | HumanEval | StructEval | QUALITY |
| Winogrande | Openbook QA | MATH-500 | MBPP | ZeroSCROLLS |
Subset of language-specific and domain-specific benchmarks coverd by eval-framework:
| Multilingual | Specialized | Safety & Bias | Efficiency Metrics |
|---|---|---|---|
| WMT Translation | MMLU | TruthfulQA | Compression ratios |
| FLORES-200 | Legal (CaseHold) | Winogender | Runtime |
| Multilingual MMLU | Scientific (SciQ) | ||
| German/Finnish tasks |
Tasks focused on logical reasoning, text distillation, instruction following, and output control. Examples include:
Tasks emphasizing classification, reasoning, and open QA. Examples include:
Tasks designed for long-context scenarios, including QA, summarization, and aggregation. Examples include:
Evaluation metrics include:
For the full list of tasks and metrics, see Detailed Task Table.
Eval-Framework provides a unified interface for evaluating language models across diverse benchmarks. The framework follows this interaction model:
BaseLLM interface (HuggingFace, OpenAI, custom APIs)BaseTask (completion, loglikelihood, or LLM-judge based)BaseMetric classespip install eval_framework[transformers]
from functools import partial
from pathlib import Path
from eval_framework.llm.huggingface import HFLLM
from eval_framework.main import main
from eval_framework.tasks.eval_config import EvalConfig
from template_formatting.formatter import HFFormatter
# Define your model
class MyHuggingFaceModel(HFLLM):
LLM_NAME = "microsoft/DialoGPT-medium"
DEFAULT_FORMATTER = partial(HFFormatter, "microsoft/DialoGPT-medium")
if __name__ == "__main__":
# Initialize your model
llm = MyHuggingFaceModel()
# Running evaluation on MMLU abstract algebra task using 5 few-shot examples and 10 samples
config = EvalConfig(
output_dir=Path("./eval_results"),
num_fewshot=5,
num_samples=10,
task_name="MMLU",
task_subjects=["abstract_algebra", "astronomy"],
llm_class=MyHuggingFaceModel,
)
# Run evaluation and get results
results = main(llm=llm, config=config)
./eval_results/ for detailed outputs and use our results guide to interpret themIf you use eval-framework in your research, please cite:
@software{eval_framework,
author={Aleph Alpha Research},
title={Aleph Alpha Eval Framework},
year={2026},
version = {x.y.z},
url={https://github.com/Aleph-Alpha-Research/eval-framework}
}
This project is licensed under the Apache License 2.0.
The constituent tasks' datasets can be subject to specific, more restrictive license terms by third parties. By using eval-framework you agree to adhere to and be bound by such license terms in the individual case.
This project has received funding from the European Union’s Digital Europe Programme under grant agreement No. 101195233 (OpenEuroLLM).
The contents of this publication are the sole responsibility of the OpenEuroLLM consortium and do not necessarily reflect the opinion of the European Union.
Python
100.0%