Kevin355/Who_and_When

Dataset

configs:

13

15 commits

1 linked in READMEs

updated Dec 21, 2025

See the code

README


configs:

  • config_name: Algorithm-Generated data_files: "Algorithm-Generated.parquet"
  • config_name: Hand-Crafted data_files: "Hand-Crafted.parquet"

Who&When: #1 Benchmark for MAS automated failure attribution.

  • 184 annotated failure tasks collected from
  • Fine-grained annotations for each failure, including:
    • The failure-responsible agent (who failed),
    • The decisive error step (when the critical error occurred),
    • A natural language explanation of the failure.

The dataset covers a wide range of realistic multi-agent scenarios based on queries from GAIA and AssistantBench. It serves as a foundational resource for developing and evaluating methods that aim to automatically pinpoint the causes of failures in complex agentic systems. We follow the following guide to annotate these failure logs. More information could be found in the paper.

Reference

If you find it useful, please consider citing our work:

@article{zhang2025agent,
  title={Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems},
  author={Zhang, Shaokun and Yin, Ming and Zhang, Jieyu and Liu, Jiale and Han, Zhiguang and Zhang, Jingyang and Li, Beibin and Wang, Chi and Wang, Huazheng and Chen, Yiran and others},
  journal={arXiv preprint arXiv:2505.00212},
  year={2025}
}

Contributors

Kevin355

15 commits

Kevin355/Who_and_When

Dataset

configs:

13

15 commits

1 linked in READMEs

updated Dec 21, 2025

See the code

README


configs:

  • config_name: Algorithm-Generated data_files: "Algorithm-Generated.parquet"
  • config_name: Hand-Crafted data_files: "Hand-Crafted.parquet"

Who&When: #1 Benchmark for MAS automated failure attribution.

  • 184 annotated failure tasks collected from
  • Fine-grained annotations for each failure, including:
    • The failure-responsible agent (who failed),
    • The decisive error step (when the critical error occurred),
    • A natural language explanation of the failure.

The dataset covers a wide range of realistic multi-agent scenarios based on queries from GAIA and AssistantBench. It serves as a foundational resource for developing and evaluating methods that aim to automatically pinpoint the causes of failures in complex agentic systems. We follow the following guide to annotate these failure logs. More information could be found in the paper.

Reference

If you find it useful, please consider citing our work:

@article{zhang2025agent,
  title={Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems},
  author={Zhang, Shaokun and Yin, Ming and Zhang, Jieyu and Liu, Jiale and Han, Zhiguang and Zhang, Jingyang and Li, Beibin and Wang, Chi and Wang, Huazheng and Chen, Yiran and others},
  journal={arXiv preprint arXiv:2505.00212},
  year={2025}
}

Contributors

Kevin355

15 commits