Public results and task definitions for FrontierHarness Eval
JavaScript
267
44 commits
updated Sep 8, 2026
frontierharness.org
30 commits
14 commits
yangzhang33/eval_benchmark
0
runta-dev/frontier-harness-eval
pan-webis-de/pan-code
Code used for evaluation and baselines in the PAN shared tasks.
47
datacurve/deep-swe
76
RollingRo11/evaltron-experiments
Interp Experiments on the Eval Aware Finetune of Llama-Nemotron-49B!
yale-nlp/MMVU-evaluation-results
claw-eval/claw-eval
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
773
bigcode-project/bigcode-evaluation-harness
A framework for the evaluation of autoregressive code generation language models.
1,062
57.2%
Shell
32.6%
Python
10.2%