OpenAgentSafety (OAS) is an open-source benchmark built on top of TheAgentCompany to systematically evaluate the safety of LLM-based agents operating in realistic, high-risk environments. Agents interact with real tools like file systems, terminals, browsers, and messaging platforms, and must navigate complex multi-turn tasks involving ambiguous, conflicting, or adversarial user instructions. OAS tasks are grounded in practical deployment scenarios and designed to reveal safety failures that occur only during dynamic multi-step interactions.
This dataset contains tasks designed for evaluating agent behavior under adversarial or ambiguous conditions.
Explanations:
workspace/: input files required by the task (e.g., .csv, .txt)Citation coming soon. For now, please refer to the GitHub repository:
All content in this dataset repository is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0).
We welcome contributions and issue reports!
13 commits
OpenAgentSafety (OAS) is an open-source benchmark built on top of TheAgentCompany to systematically evaluate the safety of LLM-based agents operating in realistic, high-risk environments. Agents interact with real tools like file systems, terminals, browsers, and messaging platforms, and must navigate complex multi-turn tasks involving ambiguous, conflicting, or adversarial user instructions. OAS tasks are grounded in practical deployment scenarios and designed to reveal safety failures that occur only during dynamic multi-step interactions.
This dataset contains tasks designed for evaluating agent behavior under adversarial or ambiguous conditions.
Explanations:
workspace/: input files required by the task (e.g., .csv, .txt)Citation coming soon. For now, please refer to the GitHub repository:
All content in this dataset repository is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0).
We welcome contributions and issue reports!
13 commits