600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.
The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.
import json
from datasets import load_dataset
runs = load_dataset(
"secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]
trace_jsonl contains the complete original JSONL trace as a string, preserving
provider-specific event fields. grades is a list of criterion IDs, titles,
rubric text (match_criteria), verdicts and judge explanations (reasoning).
The explanations describe grading decisions; they are separate from agent events.
| Fields | Contents |
|---|---|
run_id, model, trial, task_id, difficulty | Stable run identity and filters |
task_title, instructions, response | Agent request and original answer |
all_pass, criteria_passed, criteria_total, criterion_pass_rate, grades | Final grading |
trace_jsonl, trace_event_count, agent_turns, tool_calls | Full trace and execution counts |
*_path, *_sha256, judge_model | Original artifacts and provenance |
To download all original artifacts:
hf download secondstate/finance-agents-benchmark-traces \
--repo-type dataset --revision v1.0 --local-dir ./fab-traces
runs/<model>/trial-<NN>/<task>/ contains response.md, final scores.json,
original run.json and task.json, plus transcript.jsonl.gz and
metrics.json.gz. Use gzip -dc <file.gz> to read compressed files.
Final grading rubrics are in grading-tasks/; aggregate results and usage are in
reports. manifest.json records file hashes.
The judge is gpt-6-luna at maximum reasoning. A task passes only if every
criterion passes. Grades come from the full regrade completed on 27 September
2026, covering all 600 answers; these are existing evaluations, not new runs.
| Model | Task pass rate | Criterion pass rate |
|---|---|---|
| DeepSeek V4.1 Flash | 60.0% | 81.0% |
| GPT-6 Sol | 58.7% | 83.4% |
| GPT-6 Luna | 50.7% | 79.5% |
| GLM 5.3 Flash | 47.3% | 76.2% |
Original run snapshots and final grading rubrics are both retained. Agent
instructions are unchanged; some rubric criteria were clarified before the
full regrade. Grading criteria were kept outside the agent workspace.
Answers, metadata and traces preserve their original bytes; final score files
omit the local source_run path. Failed attempts, scratch files and intermediate
grading checkpoints are excluded. Results describe one synthetic company.
Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.
600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.
The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.
import json
from datasets import load_dataset
runs = load_dataset(
"secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]
trace_jsonl contains the complete original JSONL trace as a string, preserving
provider-specific event fields. grades is a list of criterion IDs, titles,
rubric text (match_criteria), verdicts and judge explanations (reasoning).
The explanations describe grading decisions; they are separate from agent events.
| Fields | Contents |
|---|---|
run_id, model, trial, task_id, difficulty | Stable run identity and filters |
task_title, instructions, response | Agent request and original answer |
all_pass, criteria_passed, criteria_total, criterion_pass_rate, grades | Final grading |
trace_jsonl, trace_event_count, agent_turns, tool_calls | Full trace and execution counts |
*_path, *_sha256, judge_model | Original artifacts and provenance |
To download all original artifacts:
hf download secondstate/finance-agents-benchmark-traces \
--repo-type dataset --revision v1.0 --local-dir ./fab-traces
runs/<model>/trial-<NN>/<task>/ contains response.md, final scores.json,
original run.json and task.json, plus transcript.jsonl.gz and
metrics.json.gz. Use gzip -dc <file.gz> to read compressed files.
Final grading rubrics are in grading-tasks/; aggregate results and usage are in
reports. manifest.json records file hashes.
The judge is gpt-6-luna at maximum reasoning. A task passes only if every
criterion passes. Grades come from the full regrade completed on 27 September
2026, covering all 600 answers; these are existing evaluations, not new runs.
| Model | Task pass rate | Criterion pass rate |
|---|---|---|
| DeepSeek V4.1 Flash | 60.0% | 81.0% |
| GPT-6 Sol | 58.7% | 83.4% |
| GPT-6 Luna | 50.7% | 79.5% |
| GLM 5.3 Flash | 47.3% | 76.2% |
Original run snapshots and final grading rubrics are both retained. Agent
instructions are unchanged; some rubric criteria were clarified before the
full regrade. Grading criteria were kept outside the agent workspace.
Answers, metadata and traces preserve their original bytes; final score files
omit the local source_run path. Failed attempts, scratch files and intermediate
grading checkpoints are excluded. Results describe one synthetic company.
Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.