Paper: Fara-7B: An Efficient Agentic Model for Computer Use
Universal Verifier: The Art of Building Verifiers for Computer Use Agents
test_v2 split (alongside the original test split, which is preserved unchanged). The test_v2 split is listed first and is the dataset-viewer default. A side-by-side per-task V1↔V2 diff (task summaries and rubric JSON) is hosted on the Fara repo: WebTailBench V1↔V2 diff.WebTailBench is a comprehensive evaluation benchmark designed to assess Computer-Using Agent (CUA) models' performance across diverse, realistic web-based tasks. This dataset was first released as part of our Fara tech report and is formally presented in the paper Fara-7B: An Efficient Agentic Model for Computer Use. The benchmark consists of:
WebTailBench addresses critical gaps in existing benchmarks by providing:
test (V1, original Nov-2025 cutoff): WebTailBench-v1-rubrics.tsvtest_v2 (V2, May-2026 refresh — recommended): WebTailBench-v2-rubrics.tsvBoth splits share the same 609 (benchmark, id) pairs. See the
V1↔V2 diff page
for per-task changes.
Data Fields (both splits):
benchmark: Task category (e.g., "flights", "hotels", "shopping", "restaurants", "activities", "ticketing", "real-estate", "jobs", "shopping_list", "comparison_shopping", "compositional_tasks")id: Unique task identifiertask_summary: Detailed task description with specific requirements and fallback instructionsprecomputed_rubric: JSON-encoded rubric used by the Universal Verifier (criteria, descriptions, max_points)Task Categories:
Single-skill tasks (8 categories):
Multi-step tasks (3 categories): 9. Shopping List (51 tasks) - Adding multiple items to cart 10. Comparison Shopping (57 tasks) - Cross-site price comparisons 11. Compositional Tasks (55 tasks) - Complex multi-step operations
Total: 609 tasks
Data Fields:
TaskID: Unique identifier for the harmful taskTaskToBeRefused: Description of the harmful task that should be refusedHarmful Task Categories (111 total tasks across 7 categories):
WebTailBench was created to address significant limitations in existing web agent benchmarks:
Tasks are derived from high-traffic commercial websites across multiple domains, reflecting actual human information needs and behaviors. All 609 tasks in the main benchmark were hand-verified by human annotators to ensure achievability.
Tasks are designed to remain valid through November 2025, after which periodic refreshes may occur. Some categories are particularly time-sensitive:
Breakdown of WebTailBench results for each of its 11 segments. Averages over three independent runs, penalizing any tasks which did not finish. The first 8 segments test a single skill or objective usually on a single website, the remaining three are more difficult multi-step or cross-site tasks.
| WebTailBench | Num Tasks | SoM 4.5 | SoM o3 | SoM 4o | GLM-4.1V 9B-Thinking | OAI Comp. Use-Prev | UI-TARS 1.5-7B | Fara 7B |
|---|---|---|---|---|---|---|---|---|
| SoM Agents | Computer Use Models | |||||||
| Shopping | 56 | 62.5 | 71.4 | 38.1 | 31.0 | 42.3 | 41.1 | 52.4 |
| Flights | 51 | 60.1 | 39.2 | 11.1 | 10.5 | 17.6 | 10.5 | 37.9 |
| Hotels | 52 | 68.6 | 56.4 | 31.4 | 19.9 | 26.9 | 35.3 | 53.8 |
| Restaurants | 52 | 67.9 | 59.6 | 47.4 | 32.1 | 35.9 | 22.4 | 47.4 |
| Activities | 80 | 70.4 | 62.9 | 41.7 | 26.3 | 30.4 | 9.6 | 36.3 |
| Ticketing | 57 | 58.5 | 56.7 | 37.4 | 35.7 | 49.7 | 30.4 | 38.6 |
| Real-Estate | 48 | 34.0 | 17.4 | 20.1 | 16.0 | 9.0 | 9.7 | 23.6 |
| Jobs/Careers | 50 | 49.3 | 44.0 | 32.7 | 22.7 | 20.7 | 20.7 | 28.0 |
| Shopping List (2 items) | 51 | 66.0 | 62.7 | 17.0 | 7.8 | 34.0 | 20.9 | 49.0 |
| Comparison Shopping | 57 | 67.3 | 59.1 | 27.5 | 22.8 | 1.2 | 8.8 | 32.7 |
| Compositional Tasks | 55 | 51.5 | 39.4 | 26.7 | 17.0 | 10.3 | 9.1 | 23.0 |
| Macro Avg. | 609 | 59.7 | 51.7 | 30.1 | 22.0 | 25.3 | 19.9 | 38.4 |
| Micro Avg. | 609 | 60.4 | 52.7 | 30.8 | 22.4 | 25.7 | 19.5 | 38.4 |
Performance varies significantly across categories, with models generally performing better on:
Per-task WebTailBench statistics for different models. All metrics are reported per task.
| Model | Cost ($) per Task | Accuracy | Actions per Task | Input Tok per Task | Output Tok per Task |
|---|---|---|---|---|---|
| SoM Agents | |||||
| SoM Agent (4.5) | 0.595 | 60.4 | 29.8 ± 26.6 | 279k ± 343k | 17.6k ± 26.0k |
| SoM Agent (o3) | 0.948 | 53.0 | 41.1 ± 34.2 | 390k ± 405k | 20.9k ± 23.4k |
| SoM Agent (4o) | 0.418 | 30.0 | 18.4 ± 18.8 | 157k ± 237k | 2.6k ± 2.6k |
| GLM-4.1V 9B-Thinking | 0.044 | 22.4 | 23.8 ± 27.9 | 117k ± 153k | 12.8k ± 15.6k |
| Computer Use Models | |||||
| OAI Comp. Use-Prev | 1.523 | 25.7 | 58.8 ± 35.4 | 493k ± 355k | 3.6k ± 2.2k |
| UI-TARS 1.5-7B | 0.133 | 19.5 | 41.1 ± 32.4 | 659k ± 631k | 3.4k ± 2.9k |
| Fara 7B | 0.069 | 38.4 | 41.1 ± 33.1 | 343k ± 323k | 2.4k ± 1.9k |
WebTailBench is designed for assessing breadth of skills and mastery of deeply chained tasks:
Positive impacts:
Potential concerns: We advise running these evaluations in a sandboxed environment without access to sensitive or personal information (e.g. a credit card or delivery address) so that real-world effects are not manifested. Risks include:
Known biases:
MIT License
If you use Fara in your research, please cite our work:
@article{Awadallah2025Fara7B,
title={Fara-7B: An Efficient Agentic Model for Computer Use},
author={Ahmed Awadallah and Yash Lara and Raghav Magazine and Hussein Mozannar and Akshay Nambi and Yash Pandya and Aravind Rajeswaran and Corby Rosset and Alexey Taymanov and Vibhav Vineet and Spencer Whitehead and Andrew Zhao},
journal={arXiv preprint arXiv:2511.19663},
year={2025},
url={https://huggingface.co/papers/2511.19663}
}
Created by Microsoft Research AI Frontiers. All tasks were hand-verified by human annotators to ensure quality and achievability.
WebTailBench includes a Task Verification system that:
For questions or issues regarding WebTailBench, please contact [contact information to be added].
Last updated: November 2025
Paper: Fara-7B: An Efficient Agentic Model for Computer Use
Universal Verifier: The Art of Building Verifiers for Computer Use Agents
test_v2 split (alongside the original test split, which is preserved unchanged). The test_v2 split is listed first and is the dataset-viewer default. A side-by-side per-task V1↔V2 diff (task summaries and rubric JSON) is hosted on the Fara repo: WebTailBench V1↔V2 diff.WebTailBench is a comprehensive evaluation benchmark designed to assess Computer-Using Agent (CUA) models' performance across diverse, realistic web-based tasks. This dataset was first released as part of our Fara tech report and is formally presented in the paper Fara-7B: An Efficient Agentic Model for Computer Use. The benchmark consists of:
WebTailBench addresses critical gaps in existing benchmarks by providing:
test (V1, original Nov-2025 cutoff): WebTailBench-v1-rubrics.tsvtest_v2 (V2, May-2026 refresh — recommended): WebTailBench-v2-rubrics.tsvBoth splits share the same 609 (benchmark, id) pairs. See the
V1↔V2 diff page
for per-task changes.
Data Fields (both splits):
benchmark: Task category (e.g., "flights", "hotels", "shopping", "restaurants", "activities", "ticketing", "real-estate", "jobs", "shopping_list", "comparison_shopping", "compositional_tasks")id: Unique task identifiertask_summary: Detailed task description with specific requirements and fallback instructionsprecomputed_rubric: JSON-encoded rubric used by the Universal Verifier (criteria, descriptions, max_points)Task Categories:
Single-skill tasks (8 categories):
Multi-step tasks (3 categories): 9. Shopping List (51 tasks) - Adding multiple items to cart 10. Comparison Shopping (57 tasks) - Cross-site price comparisons 11. Compositional Tasks (55 tasks) - Complex multi-step operations
Total: 609 tasks
Data Fields:
TaskID: Unique identifier for the harmful taskTaskToBeRefused: Description of the harmful task that should be refusedHarmful Task Categories (111 total tasks across 7 categories):
WebTailBench was created to address significant limitations in existing web agent benchmarks:
Tasks are derived from high-traffic commercial websites across multiple domains, reflecting actual human information needs and behaviors. All 609 tasks in the main benchmark were hand-verified by human annotators to ensure achievability.
Tasks are designed to remain valid through November 2025, after which periodic refreshes may occur. Some categories are particularly time-sensitive:
Breakdown of WebTailBench results for each of its 11 segments. Averages over three independent runs, penalizing any tasks which did not finish. The first 8 segments test a single skill or objective usually on a single website, the remaining three are more difficult multi-step or cross-site tasks.
| WebTailBench | Num Tasks | SoM 4.5 | SoM o3 | SoM 4o | GLM-4.1V 9B-Thinking | OAI Comp. Use-Prev | UI-TARS 1.5-7B | Fara 7B |
|---|---|---|---|---|---|---|---|---|
| SoM Agents | Computer Use Models | |||||||
| Shopping | 56 | 62.5 | 71.4 | 38.1 | 31.0 | 42.3 | 41.1 | 52.4 |
| Flights | 51 | 60.1 | 39.2 | 11.1 | 10.5 | 17.6 | 10.5 | 37.9 |
| Hotels | 52 | 68.6 | 56.4 | 31.4 | 19.9 | 26.9 | 35.3 | 53.8 |
| Restaurants | 52 | 67.9 | 59.6 | 47.4 | 32.1 | 35.9 | 22.4 | 47.4 |
| Activities | 80 | 70.4 | 62.9 | 41.7 | 26.3 | 30.4 | 9.6 | 36.3 |
| Ticketing | 57 | 58.5 | 56.7 | 37.4 | 35.7 | 49.7 | 30.4 | 38.6 |
| Real-Estate | 48 | 34.0 | 17.4 | 20.1 | 16.0 | 9.0 | 9.7 | 23.6 |
| Jobs/Careers | 50 | 49.3 | 44.0 | 32.7 | 22.7 | 20.7 | 20.7 | 28.0 |
| Shopping List (2 items) | 51 | 66.0 | 62.7 | 17.0 | 7.8 | 34.0 | 20.9 | 49.0 |
| Comparison Shopping | 57 | 67.3 | 59.1 | 27.5 | 22.8 | 1.2 | 8.8 | 32.7 |
| Compositional Tasks | 55 | 51.5 | 39.4 | 26.7 | 17.0 | 10.3 | 9.1 | 23.0 |
| Macro Avg. | 609 | 59.7 | 51.7 | 30.1 | 22.0 | 25.3 | 19.9 | 38.4 |
| Micro Avg. | 609 | 60.4 | 52.7 | 30.8 | 22.4 | 25.7 | 19.5 | 38.4 |
Performance varies significantly across categories, with models generally performing better on:
Per-task WebTailBench statistics for different models. All metrics are reported per task.
| Model | Cost ($) per Task | Accuracy | Actions per Task | Input Tok per Task | Output Tok per Task |
|---|---|---|---|---|---|
| SoM Agents | |||||
| SoM Agent (4.5) | 0.595 | 60.4 | 29.8 ± 26.6 | 279k ± 343k | 17.6k ± 26.0k |
| SoM Agent (o3) | 0.948 | 53.0 | 41.1 ± 34.2 | 390k ± 405k | 20.9k ± 23.4k |
| SoM Agent (4o) | 0.418 | 30.0 | 18.4 ± 18.8 | 157k ± 237k | 2.6k ± 2.6k |
| GLM-4.1V 9B-Thinking | 0.044 | 22.4 | 23.8 ± 27.9 | 117k ± 153k | 12.8k ± 15.6k |
| Computer Use Models | |||||
| OAI Comp. Use-Prev | 1.523 | 25.7 | 58.8 ± 35.4 | 493k ± 355k | 3.6k ± 2.2k |
| UI-TARS 1.5-7B | 0.133 | 19.5 | 41.1 ± 32.4 | 659k ± 631k | 3.4k ± 2.9k |
| Fara 7B | 0.069 | 38.4 | 41.1 ± 33.1 | 343k ± 323k | 2.4k ± 1.9k |
WebTailBench is designed for assessing breadth of skills and mastery of deeply chained tasks:
Positive impacts:
Potential concerns: We advise running these evaluations in a sandboxed environment without access to sensitive or personal information (e.g. a credit card or delivery address) so that real-world effects are not manifested. Risks include:
Known biases:
MIT License
If you use Fara in your research, please cite our work:
@article{Awadallah2025Fara7B,
title={Fara-7B: An Efficient Agentic Model for Computer Use},
author={Ahmed Awadallah and Yash Lara and Raghav Magazine and Hussein Mozannar and Akshay Nambi and Yash Pandya and Aravind Rajeswaran and Corby Rosset and Alexey Taymanov and Vibhav Vineet and Spencer Whitehead and Andrew Zhao},
journal={arXiv preprint arXiv:2511.19663},
year={2025},
url={https://huggingface.co/papers/2511.19663}
}
Created by Microsoft Research AI Frontiers. All tasks were hand-verified by human annotators to ensure quality and achievability.
WebTailBench includes a Task Verification system that:
For questions or issues regarding WebTailBench, please contact [contact information to be added].
Last updated: November 2025