secondstate/finance-agents-benchmark-traces

Dataset

FAB — Agent Traces and Grading

1

12 commits

1 linked in READMEs

updated Sep 27, 2026

See the code

README

FAB — Agent Traces and Grading

600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.

The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.

Use

import json
from datasets import load_dataset

runs = load_dataset(
    "secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]

trace_jsonl contains the complete original JSONL trace as a string, preserving provider-specific event fields. grades is a list of criterion IDs, titles, rubric text (match_criteria), verdicts and judge explanations (reasoning). The explanations describe grading decisions; they are separate from agent events.

FieldsContents
run_id, model, trial, task_id, difficultyStable run identity and filters
task_title, instructions, responseAgent request and original answer
all_pass, criteria_passed, criteria_total, criterion_pass_rate, gradesFinal grading
trace_jsonl, trace_event_count, agent_turns, tool_callsFull trace and execution counts
*_path, *_sha256, judge_modelOriginal artifacts and provenance

To download all original artifacts:

hf download secondstate/finance-agents-benchmark-traces \
  --repo-type dataset --revision v1.0 --local-dir ./fab-traces

runs/<model>/trial-<NN>/<task>/ contains response.md, final scores.json, original run.json and task.json, plus transcript.jsonl.gz and metrics.json.gz. Use gzip -dc <file.gz> to read compressed files. Final grading rubrics are in grading-tasks/; aggregate results and usage are in reports. manifest.json records file hashes.

Grading and scope

The judge is gpt-6-luna at maximum reasoning. A task passes only if every criterion passes. Grades come from the full regrade completed on 27 September 2026, covering all 600 answers; these are existing evaluations, not new runs.

ModelTask pass rateCriterion pass rate
DeepSeek V4.1 Flash60.0%81.0%
GPT-6 Sol58.7%83.4%
GPT-6 Luna50.7%79.5%
GLM 5.3 Flash47.3%76.2%

Original run snapshots and final grading rubrics are both retained. Agent instructions are unchanged; some rubric criteria were clarified before the full regrade. Grading criteria were kept outside the agent workspace. Answers, metadata and traces preserve their original bytes; final score files omit the local source_run path. Failed attempts, scratch files and intermediate grading checkpoints are excluded. Results describe one synthetic company.

Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.

agents
agent-traces
benchmark
evaluation
finance
financial-due-diligence
synthetic
tool-use

secondstate/finance-agents-benchmark-traces

Dataset

FAB — Agent Traces and Grading

1

12 commits

1 linked in READMEs

updated Sep 27, 2026

See the code

README

FAB — Agent Traces and Grading

600 completed agent runs: four models × 50 tasks × three trials. Agents investigate a synthetic company's data room and answer financial due-diligence questions. Each row pairs a full execution trace with the task, final answer, grading criteria, pass/fail verdicts and judge explanations. Tasks 041 and 049 are included.

The benchmark dataset contains the shared data room and tasks. The GitHub repository contains the execution and grading harness.

Use

import json
from datasets import load_dataset

runs = load_dataset(
    "secondstate/finance-agents-benchmark-traces", revision="v1.0", split="test"
)
run = runs[0]
events = [json.loads(line) for line in run["trace_jsonl"].splitlines()]
grades = run["grades"]

trace_jsonl contains the complete original JSONL trace as a string, preserving provider-specific event fields. grades is a list of criterion IDs, titles, rubric text (match_criteria), verdicts and judge explanations (reasoning). The explanations describe grading decisions; they are separate from agent events.

FieldsContents
run_id, model, trial, task_id, difficultyStable run identity and filters
task_title, instructions, responseAgent request and original answer
all_pass, criteria_passed, criteria_total, criterion_pass_rate, gradesFinal grading
trace_jsonl, trace_event_count, agent_turns, tool_callsFull trace and execution counts
*_path, *_sha256, judge_modelOriginal artifacts and provenance

To download all original artifacts:

hf download secondstate/finance-agents-benchmark-traces \
  --repo-type dataset --revision v1.0 --local-dir ./fab-traces

runs/<model>/trial-<NN>/<task>/ contains response.md, final scores.json, original run.json and task.json, plus transcript.jsonl.gz and metrics.json.gz. Use gzip -dc <file.gz> to read compressed files. Final grading rubrics are in grading-tasks/; aggregate results and usage are in reports. manifest.json records file hashes.

Grading and scope

The judge is gpt-6-luna at maximum reasoning. A task passes only if every criterion passes. Grades come from the full regrade completed on 27 September 2026, covering all 600 answers; these are existing evaluations, not new runs.

ModelTask pass rateCriterion pass rate
DeepSeek V4.1 Flash60.0%81.0%
GPT-6 Sol58.7%83.4%
GPT-6 Luna50.7%79.5%
GLM 5.3 Flash47.3%76.2%

Original run snapshots and final grading rubrics are both retained. Agent instructions are unchanged; some rubric criteria were clarified before the full regrade. Grading criteria were kept outside the agent workspace. Answers, metadata and traces preserve their original bytes; final score files omit the local source_run path. Failed attempts, scratch files and intermediate grading checkpoints are excluded. Results describe one synthetic company.

Source: published runs at commit 4c78c7c. License: CC BY 4.0. Cite SecondState's Finance Agents Benchmark and the dataset revision when reusing these results.

agents
agent-traces
benchmark
evaluation
finance
financial-due-diligence
synthetic
tool-use