Reproducible agent-evaluation experiment: rare failures, selection errors, and regret under fixed budgets
Python
0
3 commits
updated Oct 5, 2026
An agent evaluator can make more wrong picks and still lose less utility.
A reproducible Python experiment on rare failures, fixed evaluation budgets, and the gap between selection accuracy and the cost of a mistake.
Results · Full comparison · Protocol · Raw data
Same budget. Same tasks. Same 1,024 simulated worlds.
| At 25,600 draws per policy per world | Uniform, full budget | Variance allocation |
|---|---|---|
| Concentrated risk: wrong picks | 455 / 1,024 | 539 / 1,024 |
| Concentrated risk: mean utility lost | 0.0268 | 0.0193 |
| Diffuse risk: wrong picks | 541 / 1,024 | 602 / 1,024 |
| Diffuse risk: mean utility lost | 0.0208 | 0.0229 |
Selection accuracy and decision cost can rank evaluators differently. In the concentrated setting, variance allocation increased the wrong-selection rate from 44.4% to 52.6%, while reducing mean utility lost by 27.8% (0.0268 to 0.0193). It picked the true best agent less often, but its mistakes were less costly on average. Under diffuse risk, both metrics were worse.
For an evaluation pipeline, the practical lesson is to measure both how often it selects a suboptimal agent and how much expected utility that choice loses. Accuracy counts a near tie and a costly mistake equally. Regret captures the utility gap, but does not enforce a catastrophic-risk limit. These synthetic results motivate reporting both; they do not establish deployment safety or a general advantage for adaptive evaluation.
Python 3.13. CPU only. No model API, credentials or private data.
git clone https://github.com/itsloganmann/rare-failure-eval.git
cd rare-failure-eval
python3.13 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python run_experiment.py --mode smoke --out outputs/smoke
Read outputs/smoke/REPORT.md to inspect the smoke run. Use a new output directory
for each run. The published findings come from the full run, not smoke mode.
.venv/bin/python -m pytest -q
.venv/bin/python run_experiment.py --mode primary --approved-primary --out outputs/replication
.venv/bin/python run_experiment.py --verify --out outputs/replication
.venv/bin/python independent_audit.py outputs/replication
.venv/bin/python make_figures.py outputs/replication/summary.json --out outputs/figures
The primary flag is an execution gate, not evidence of peer review. To verify the distributed run without resampling:
mkdir -p outputs/published
cp results/*.json results/REPORT.md outputs/published/
gzip -dc results/results.jsonl.gz > outputs/published/results.jsonl
.venv/bin/python run_experiment.py --verify --out outputs/published
.venv/bin/python independent_audit.py outputs/published
The manifest hashes the frozen study code and outputs. The separate arithmetic audit reconstructs weighted estimates, rankings, ties, selection error and regret from saved sufficient statistics.
Four synthetic agents, eight tasks, three risk settings, two budgets and four
sampling policies. Utility is reward - 800 * failure, using fixed task weights.
Regret is the true best agent's expected utility minus the selected agent's.
| Policy | Allocation |
|---|---|
| Uniform, full budget | Every draw goes into balanced estimation |
| Uniform, pilot matched | A pilot followed by balanced independent estimation |
| Variance allocation | A four-draw-per-cell pilot sets a frozen Neyman allocation, with a uniform floor |
| Follow pilot failures | Allocate toward observed failures; otherwise use uniform sampling |
The full run contains 24,576 records and 393,216,000 simulated draws. All conditions · Paired intervals · Audit
Inspired by Active Evaluation of General Agents, Lanctot et al. (2026). This independent experiment studies rare-loss utility and sampling allocation; it does not reproduce that paper's Elo or Soft Condorcet algorithms. No employer data or institutional endorsement is involved.
Reproducible agent-evaluation experiment: rare failures, selection errors, and regret under fixed budgets
Python
0
3 commits
updated Oct 5, 2026
An agent evaluator can make more wrong picks and still lose less utility.
A reproducible Python experiment on rare failures, fixed evaluation budgets, and the gap between selection accuracy and the cost of a mistake.
Results · Full comparison · Protocol · Raw data
Same budget. Same tasks. Same 1,024 simulated worlds.
| At 25,600 draws per policy per world | Uniform, full budget | Variance allocation |
|---|---|---|
| Concentrated risk: wrong picks | 455 / 1,024 | 539 / 1,024 |
| Concentrated risk: mean utility lost | 0.0268 | 0.0193 |
| Diffuse risk: wrong picks | 541 / 1,024 | 602 / 1,024 |
| Diffuse risk: mean utility lost | 0.0208 | 0.0229 |
Selection accuracy and decision cost can rank evaluators differently. In the concentrated setting, variance allocation increased the wrong-selection rate from 44.4% to 52.6%, while reducing mean utility lost by 27.8% (0.0268 to 0.0193). It picked the true best agent less often, but its mistakes were less costly on average. Under diffuse risk, both metrics were worse.
For an evaluation pipeline, the practical lesson is to measure both how often it selects a suboptimal agent and how much expected utility that choice loses. Accuracy counts a near tie and a costly mistake equally. Regret captures the utility gap, but does not enforce a catastrophic-risk limit. These synthetic results motivate reporting both; they do not establish deployment safety or a general advantage for adaptive evaluation.
Python 3.13. CPU only. No model API, credentials or private data.
git clone https://github.com/itsloganmann/rare-failure-eval.git
cd rare-failure-eval
python3.13 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python run_experiment.py --mode smoke --out outputs/smoke
Read outputs/smoke/REPORT.md to inspect the smoke run. Use a new output directory
for each run. The published findings come from the full run, not smoke mode.
.venv/bin/python -m pytest -q
.venv/bin/python run_experiment.py --mode primary --approved-primary --out outputs/replication
.venv/bin/python run_experiment.py --verify --out outputs/replication
.venv/bin/python independent_audit.py outputs/replication
.venv/bin/python make_figures.py outputs/replication/summary.json --out outputs/figures
The primary flag is an execution gate, not evidence of peer review. To verify the distributed run without resampling:
mkdir -p outputs/published
cp results/*.json results/REPORT.md outputs/published/
gzip -dc results/results.jsonl.gz > outputs/published/results.jsonl
.venv/bin/python run_experiment.py --verify --out outputs/published
.venv/bin/python independent_audit.py outputs/published
The manifest hashes the frozen study code and outputs. The separate arithmetic audit reconstructs weighted estimates, rankings, ties, selection error and regret from saved sufficient statistics.
Four synthetic agents, eight tasks, three risk settings, two budgets and four
sampling policies. Utility is reward - 800 * failure, using fixed task weights.
Regret is the true best agent's expected utility minus the selected agent's.
| Policy | Allocation |
|---|---|
| Uniform, full budget | Every draw goes into balanced estimation |
| Uniform, pilot matched | A pilot followed by balanced independent estimation |
| Variance allocation | A four-draw-per-cell pilot sets a frozen Neyman allocation, with a uniform floor |
| Follow pilot failures | Allocate toward observed failures; otherwise use uniform sampling |
The full run contains 24,576 records and 393,216,000 simulated draws. All conditions · Paired intervals · Audit
Inspired by Active Evaluation of General Agents, Lanctot et al. (2026). This independent experiment studies rare-loss utility and sampling allocation; it does not reproduce that paper's Elo or Soft Condorcet algorithms. No employer data or institutional endorsement is involved.