AIgenteur/ClawEval

49

stars

0

commits

Python

primary language

Sep 10, 2026

updated

README

🦞 ClawEval — Can Your Local or Cloud LLM Actually Do the Job?

AIgenteur

The only deterministic benchmark that tests LLMs as real AI agents — not trivia, not chat, not vibes. 59 specialized roles. Hard, verifiable tasks. Every score is reproducible. Built for OpenClaw, Hermes Agent, and any model-agnostic agent framework.

🔥 New models added regularly. Star ⭐ this repo to get notified when new results drop.

📬 Subscribe to AIgenteur Dispatch — benchmark results, local AI guides, and agent workflows in your inbox.

🎓 Join AIgenteur Launchpad (free) — community for freelancers and entrepreneurs putting AI agents to work. Want a model tested? Post your request in the community.

🤖 Sister project: AIgenteur — open-source AI cofounder for solopreneurs. ClawEval grades the agents it dispatches.

💡 You'd be surprised how well some small, fast models handle sub-agent tasks — even on an RTX 3090 (we got ours for $799). Before you rent cloud GPUs, check the scores below.

Why ClawEval?

Most benchmarks tell you a model is "smart." ClawEval tells you if it can do the work — route tickets, review code, analyze financials, draft legal docs, plan sprints, and 54 more agent roles. Each test has an exact expected answer. No LLM-as-judge. No vibes.

Built for agent frameworks like OpenClaw and Hermes Agent — any system where you need to pick the right LLM backend for specialized sub-agent roles. If your framework is model-agnostic, ClawEval shows you exactly which model to plug in for each job.

⚡ NEW: Token Efficiency Leaderboard

🏆 Which models waste the fewest tokens? →

Two models both score 10/10 — but one used 500 tokens, the other 5,000. The efficient one saves 10× on compute, cost, and latency. We rank 35 models by tokens-per-point. Think modes use 10–20× more tokens for marginal gains. This metric is hardware-independent.

🚀 NEW: ClawEval v2 — 59 Agents, 1,220 Checkpoints, Zero Ceiling Effects

🏆 The definitive dense evaluation →

Phase F scored x/10 — everyone got 8+. ClawEval v2 upgrades all 59 agents to 15–30 checkpoints each with adversarial traps, sarcasm, and near-truths. Real separation, real rankings. First results (Kimi K2.6, GLM-5.1) dropping now.

🖥️ LOCAL Models — Run on YOUR Hardware

📄 OpenClaw Backend: Local on 3090 — The Complete Guide (PDF) — An interactive walkthrough of running the full OpenClaw backend locally on an RTX 3090. Covers setup, model selection, and performance tips with step-by-step calls to action.

We test quantized open-source models that fit on hardware you already own. Find out which model is best for each agent role before you commit your VRAM. We're also testing smaller models for max context and usability — even a 16GB GPU can run capable sub-agents.

🔴 8GB VRAM🟠 12GB VRAM🟡 16GB VRAM🟢 24GB VRAM🔵 64–96GB VRAM
Qwen3.5-0.8B Q4_K_MQwen3.5-9B Q4_K_MQwen3.5-9B Q4_K_MQwen3.6-35B-A3B Q4_K_MQwen3.5-122B-A10B NVFP4
Qwen3.5-2B Q4_K_MQwen3.5-35B-A3B Q4_K_MGPT-OSS-120B GGUF
Qwen3.5-4B Q4_K_M (250K ctx)Qwen3.5-27B Q4_K_M
✅ Tested✅ Testedllama.cppllama.cpp · SGLang · vLLMSGLang · vLLM · llama.cpp

📖 VRAM Guides: 8–16GB Small Models · 16GB · 24GB · 32GB · 48GB · 64GB · 96GB — Which models fit, context limits, speed estimates

🧩 Which Tested Models Fit on Your Hardware?

No discrete GPU? Unified-memory devices can run these models directly. Estimates are conservative — they account for OS overhead (~4–8 GB) and KV cache for reasonable context windows.

Model (as tested)WeightsMac Mini M4 (16 GB)Mac Mini M4 Pro (24 GB)Mac Mini M4 Pro (48 GB)Strix Halo (96 GB GPU)DGX Spark (96 GB GPU)
Qwen3.5-0.8B Q4_K_M~0.5 GB
Qwen3.5-2B Q4_K_M~1.5 GB
Gemma-4-E2B Q4_K_M~1.5 GB
Qwen3.5-4B Q4_K_M~2.8 GB
Gemma-4-E4B Q4_K_M~2.5 GB
Ministral-3B Q4_K_M~2 GB
Ministral-8B Q4_K_M~5 GB
Qwen3.5-9B Q4_K_M~6 GB
Ministral-14B Q4_K_M~9 GB⚠️ tight
Gemma-4-26B-A4B Q4_K_M~15 GB⚠️ tight
Qwen3.5-27B Q4_K_M~17 GB⚠️ tight
Gemma-4-31B Q4_K_M~19 GB⚠️ tight
Qwen3.5-35B-A3B NVFP4~20 GB
Qwen3.6-35B-A3B Q4_K_M~20 GB
GPT-OSS-120B GGUF~61 GB
Qwen3.5-122B-A10B NVFP4~70 GB
Nemotron-3-Super-120B-A12B NVFP4~70 GB
Mistral-Small-4-119B NVFP4~70 GB

⚠️ tight = model loads but leaves little room for KV cache / long context. Short prompts only.

Strix Halo = AMD Ryzen AI Max (128 GB LPDDR5X total, up to 96 GB allocatable to GPU). DGX Spark = NVIDIA GB10 (128 GB LPDDR5X unified, ~96 GB GPU-usable). Both can run 120B-class models with room for KV cache.

Mac Mini M4 Pro 64 GB and M4 Max 128 GB configs exist but are not listed — they slot between the columns above.

☁️ CLOUD Models — Run via API

Don't have a GPU? We also test open-source models hosted on cloud providers so you can compare local vs cloud performance at every agent role. Same tests, same scoring — find out if paying for cloud is worth it, or if your local setup already matches.

ProviderModelsStatus
Alibaba Coding PlanQwen3.5-Plus — 🥇 472/590 (80%) Phase F, 86/110 (78%) Phase G✅ Tested
Alibaba Coding PlanKimi K2.5 — 473/590 (80%) Phase F, 96/110 (87%) Phase G✅ Tested
Alibaba Coding PlanGLM-5 — 456/590 (77%) Phase F, 80/110 (73%) Phase G✅ Tested
Alibaba Coding PlanMiniMax-M2.5 — 455/590 (77%) Phase F, 78/110 (71%) Phase G✅ Tested
Ollama CloudMiniMax-M2.7 — 473/590 (80%) Phase F, 83/110 (75%) Phase G✅ Tested
OpenRouterStep-3.5-Flash — 438/590 (74%) Phase F, 66/110 (60%) Phase G✅ Tested
OpenRouterTrinity-Large — 444/590 (75%) Phase F, 73/110 (66%) Phase G✅ Tested
OpenRouterNemotron-3-Nano-30B — 453/590 (77%) Phase F, 70/110 (64%) Phase G✅ Tested
OpenRouterQwen3.6-Plus — 469/590 (79%) Phase F, 91/110 (83%) Phase G✅ Tested

💡 Kimi K2.5 leads Phase F at 473/590 (80%), closely followed by M2.7 and Qwen3.5-Plus. M2.7's reasoning tokens count against max_tokens, requiring 16-24k to avoid truncation.

📊 Detailed cloud comparisons:

⚡ Token Efficiency — Which Models Waste the Fewest Tokens?

Two models both score 10/10 — but one used 500 tokens, the other 5,000. The efficient one saves 10× on compute, cost, and latency. This is hardware-independent: it measures model behavior, not GPU speed.

RankModelVRAMScoreTokens/PointVerdict
🥇Trinity-Large☁️41439Ultra-lean
🥈Mistral-119B NT🔵 64GB41949Ultra-lean
🥉GLM-5 NT☁️42161Lean
4Qwen3.5-Plus NT☁️44263Lean
5Qwen3.5-35B-A3B NT🟢 24GB43276Efficient
6Kimi K2.5 Think☁️44380Efficient
...
32Qwen3.6-35B-A3B🟢 24GB430696Heavy
35Qwen3.5-35B-A3B Think🟢 24GB3951,524Extremely Heavy

💡 Key insight: Thinking models use 10–20× more tokens than NoThink variants for similar scores. For cost-sensitive agent pipelines, NoThink wins massively on efficiency.

📊 Full Token Efficiency Leaderboard (35 models) → — complete rankings, Think vs NoThink gaps, best per VRAM tier


🚀 ClawEval v2 — Full 59-Agent Dense Evaluation

The next generation of ClawEval testing. Every agent role now has a dense constraint test with 15–30 checkpoints — no more x/10 ceiling effects. This is the definitive leaderboard.

Phase F gave every model 8–10/10 on most roles. ClawEval v2 replaces that with granular percentage scoring across all 59 agents. Same roles, dramatically harder prompts, real separation between models.

MetricPhase F (v1)ClawEval v2
Tests59 roles × 10 pts59 roles × 15–30 checkpoints
Total checkpoints5901,220
Scoringx/10 (ceiling effect)Raw % (fine-grained)
Manual-only tests60 — fully automated
Discriminating powerLow (everyone scores 8+)High (real spread)

🏅 ClawEval v2 Leaderboard — Open-Weight Models

RankModelProviderScore%Perfect
🥇DeepSeek V4 Pro☁️ DeepSeek1060/122086.9%26
🥈DeepSeek V4 Flash☁️ DeepSeek1054/122086.4%23
🥉Kimi K2.5 Think☁️ Ollama1048/122085.9%24
4Kimi K2.7 Code☁️ Ollama1038/122085.1%24
5Laguna-M.1☁️ OpenRouter1034/122084.8%25
6Qwen3.5-Plus☁️ Alibaba1031/122084.5%26
7Qwen3.6-35B-A3B🖥️ Local1029/122084.3%24
8Kimi K2.6☁️ Ollama1028/122084.3%24
9Qwen3.5-122B-A10B☁️ Ollama1025/122084.0%21
10Gemma-4-31B☁️ Ollama1024/122083.9%25
11Cobuddy☁️ OpenRouter1023/122083.9%21
12Mistral-Large-3☁️ Ollama1021/122083.7%25
13GLM-5.1☁️ Ollama1020/122083.6%26
14Nemotron-3-Super Think☁️ Ollama1016/122083.3%20
15MiniMax-M2.7 Medium☁️ Ollama1014/122083.1%20
16Qwen3.6-27B🖥️ Local TQ41012/122083.0%26
17Nemotron-3-Super NoThink☁️ Ollama996/122081.6%21
18MiniMax-M3☁️ Ollama993/122081.4%25
19MiniMax-M2.7 Think☁️ Ollama993/122081.4%19
20Nemotron-3-Nano-Omni☁️ OpenRouter991/122081.2%20
21Gemma-4-E2B🖥️ Local981/122080.4%14
22GPT-OSS-120B☁️ Ollama979/122080.2%19
23Gemini-3.5-Flash☁️ Google (OpenRouter)978/122080.2%22
24Phi-4🖥️ Local Q8977/122080.1%17
25GLM-5.2☁️ Ollama957/122078.4%23
26Laguna-XS.2☁️ OpenRouter950/122077.9%20
27GLM-5 NoThink☁️ Ollama948/122077.7%25
28Nemotron-Nano-Omni🖥️ Local IQ4948/122077.7%20
29Kimi K2.5 NoThink☁️ Ollama935/122076.6%21
30Granite-4.1 30B🖥️ Local TQ4929/122076.1%15
31Granite-4.1 8B🖥️ Local Q4929/122076.1%14
32Gemma-4-31B🖥️ Local Q4927/122076.0%25
33GLM-5 Think☁️ Ollama927/122076.0%23
34Nemotron-Nano-Omni🖥️ Local Q4925/122075.8%19
35Trinity-Large-Think☁️ OpenRouter914/122074.9%21
36Nemotron-3-Nano-30B☁️ Ollama914/122074.9%19
37Ministral-3 8B☁️ Ollama906/122074.3%18
38Ministral-3 14B☁️ Ollama888/122072.8%17
39GPT-OSS-20B☁️ Ollama885/122072.5%19
40Ministral-3 8B🖥️ Local Q4884/122072.5%16
41Ministral-3 14B🖥️ Local Q4877/122071.9%18
42Gemma-4-E4B🖥️ Local867/122071.1%15
43Ministral-3 14B Think🖥️ Local Q4858/122070.3%19
44Granite-4.1 3B🖥️ Local Q4846/122069.3%12
45Ministral-3 3B☁️ Ollama844/122069.2%14
46Ministral-3 8B Think🖥️ Local Q4791/122064.8%10
47Ministral-3 3B🖥️ Local Q4760/122062.3%12
48RNJ-1-8B☁️ Ollama750/122061.5%18
49Ministral-3 3B Think🖥️ Local Q4704/122057.7%10
50Gemma-4-A4B🖥️ Local622/122051.0%10
51Qwen3.5-9B🖥️ Local543/122044.5%6
52Qwen3.5-4B🖥️ Local374/122030.7%4
53LFM2.5-350M🖥️ Local308/122025.2%2
54Qwen3.5-0.8B🖥️ Local58/12204.8%0
55Qwen3.5-2B🖥️ Local50/12204.1%0

📋 Gemini-3.5-Flash — Full 59-Agent Breakdown

Per-role performance across all 59 ClawEval v2 agents. 978/1,220 checkpoints (80.2%) — 22 roles at 100%, 12 roles below 50%.

TestAgent RoleScore%
H-01Router / Triage Agent29/3097%
H-02Input Validator / Sanitizer29/3097%
H-03Heartbeat / Health Monitor14/1593%
H-04Notification / Alert Agent21/3070%
H-05Sentiment Analysis Agent29/3097%
H-06FAQ Generation Agent15/15100%
H-07Translation Agent15/15100%
H-08Calendar / Scheduling Agent12/2060%
H-09Research / Web Search Agent30/30100%
H-10Content Writer / Blog Writer18/2090%
H-11Editor Agent29/3097%
H-12Content Planner / Strategist0/300%
H-13Email Drafting / Summarization43/4596%
H-14Document Summarization Agent15/15100%
H-15Meeting Notes / Transcription Agent35/35100%
H-16Social Media Scouting / Monitoring57/6095%
H-17Social Media Content Agent19/2095%
H-18News Aggregation Agent7/7100%
H-19Shopping / Price Comparison15/15100%
H-20Memory / Knowledge Management20/20100%
H-21RAG / Retrieval Agent4/1527%
H-22Data Analysis Agent14/1593%
H-23Website Scraping / Understanding15/15100%
H-24Image Description / Understanding20/20100%
H-25Customer Support Agent33/6055%
H-26Lead Scoring / Prospecting12/1580%
H-27Sprint / Project Summarizer15/15100%
H-28Transaction / Approval Agent19/2095%
H-29Home Automation Agent9/2045%
H-30Fitness / Health Tracking15/15100%
H-31Recipe / Cooking Agent15/15100%
H-32Personal Finance Tracking15/15100%
H-33SEO Optimization Agent4/1527%
H-34Landing Page Generator20/20100%
H-35Travel Planning Agent6/1540%
H-36Code Generation Agent30/30100%
H-37Code Review Agent15/15100%
H-38QA / Test Writing Agent13/1587%
H-39Task Planning / Decomposition1/186%
H-40Fact-Checking Agent28/3093%
H-41Critic / Review Agent20/20100%
H-42Market Research Agent14/1593%
H-43Synthesizer / Aggregator13/1587%
H-44Curriculum / Course Designer7/1547%
H-45Prototype Generator15/15100%
H-46DevOps Agent8/1553%
H-47Math / Logic Reasoning14/1593%
H-48STEM Research Analyst15/15100%
H-49Algorithm / Data Structure Explorer30/30100%
H-50Orchestrator / Manager Agent1/157%
H-51Software Architect Agent4/1527%
H-52Complex Debugger Agent13/1587%
H-53Legal Document Review6/1540%
H-54Medical / Health Analysis15/15100%
H-55Financial Analysis / Stock Research14/1593%
H-56Security Analyst Agent0/150%
H-57SRE / Incident Response13/1587%
H-58Book / Long-Form Writing19/2095%
H-59Compliance / Regulatory Agent2/1513%

📦 Evaluation Phases

ClawEval evaluates models across multiple phases of increasing difficulty:

PhaseFocusTestsScoring
A–CBasic role evaluation59 roles × system promptsQuality review
DHard prompts59 roles × adversarial promptsAutomated
EKiller tests12 precision tasksDeterministic (JSON, code exec, regex)
FRole-specific hard tests59 roles × deterministic promptsDeterministic (15+ scoring types)

Phase F — 59 Roles Across 5 Tiers

  • Tier 1 (8 roles): Router, Validator, Health Monitor, Notification, Sentiment, FAQ, Translation, Calendar
  • Tier 2 (27 roles): Research, Writer, Editor, Email, Summarization, Customer Support, Data Analysis, etc.
  • Tier 3 (11 roles): Code Gen, Code Review, QA Testing, Task Planning, Fact-Checking, Market Research, etc.
  • Tier 4 (3 roles): Math/Logic Reasoning, STEM Analysis, Algorithm Exploration
  • Tier 5 (10 roles): Orchestrator, Architect, Debugger, Legal, Medical, Financial, Security, SRE, etc.

Scoring Types

exact_json · json_numeric · keyword_detection · code_exec · constraint_check · error_count · action_items · news_dedup · recipe_scaling · compliance_issues · architecture_constraints · manual_review · and more

🚀 Quick Start

Prerequisites

python3 -m venv .venv
source .venv/bin/activate
pip install requests

Run Phase E (12 killer tests)

# Against llama.cpp server
python eval/run_phase_e.py \
  --base-url http://localhost:8080/v1 \
  --model my-model-name \
  --max-tokens 4000

# Against SGLang server (with thinking)
python eval/run_phase_e.py \
  --base-url http://192.168.1.2:8000/v1 \
  --model Qwen3.5-122B-think \
  --api-model /path/to/weights \
  --thinking-budget 16384 \
  --max-tokens 32000

# Disable thinking
python eval/run_phase_e.py \
  --base-url http://192.168.1.2:8000/v1 \
  --model Qwen3.5-122B-nothink \
  --api-model /path/to/weights \
  --nothink

Run Phase F (59 role tests)

python eval/run_phase_f.py \
  --base-url http://localhost:8080/v1 \
  --model my-model-name \
  --max-tokens 4000

# Run specific tests or tiers
python eval/run_phase_f.py ... --test-ids 1 2 3
python eval/run_phase_f.py ... --tier 3

Results

Results are saved to eval/test_results/<model-name>/phase_e/ and phase_f/ directories, including:

  • Individual response text files per test
  • phase_e_scores.json / phase_f_scores.json with full scoring breakdown

🔧 Key Features

  • OpenAI-compatible API: Works with any server exposing /v1/chat/completions (llama.cpp, SGLang, vLLM, etc.)
  • SGLang thinking control: --thinking-budget and --nothink flags for reasoning token management
  • Think-tag stripping: Automatically strips <think> tags from responses for clean scoring
  • Flexible JSON scoring: Case-insensitive keys, nested dict traversal, numeric tolerance
  • Deterministic: Every test has an exact expected answer — no LLM-as-judge

📁 Repo Structure

ClawEval/
├── eval/
│   ├── run_phase_e.py          # Phase E runner (12 killer tests)
│   ├── run_phase_f.py          # Phase F runner (59 role tests)
│   ├── phase_e_prompts.py      # Phase E test definitions
│   ├── phase_f/                # Phase F test definitions (5 tiers)
│   │   ├── __init__.py
│   │   ├── tier1.py ... tier5.py
│   ├── role_prompts.py         # 59 role system prompts
│   ├── hard_prompts.py         # Phase D adversarial prompts
│   └── test_results/           # All model evaluation results
├── docs/                       # VRAM tier guides & model selection
│   ├── 16GB, 24GB, 32GB, 48GB, 64GB, 96GB tier guides
│   └── Subagent type reference
├── RESULTS.md                  # Detailed per-role score comparison
└── README.md

🗺️ Model Testing Roadmap

This is a living benchmark. We're continuously adding new models as they release. Here's what's on the radar — and we take requests.

16GB VRAM tier (coming soon)

  • Qwen3-8B, Qwen3.5-14B
  • Phi-4 14B, Phi-4-mini 3.8B
  • Mistral Small 3.1 24B (tight fit)
  • GPT-OSS-20B MXFP4
  • Gemma 3 12B
  • Ministral 3B (ultra-fast routing/triage)

24GB VRAM tier

  • ✅ Qwen3.5-35B-A3B Q4_K_M (tested)
  • ✅ Qwen3.5-27B Q4_K_M (tested)
  • Gemma 3 27B QAT
  • More MoE models as they release

64–96GB VRAM tier

  • ✅ Qwen3.5-122B-A10B NVFP4 (tested)
  • ✅ GPT-OSS-120B GGUF · llama.cpp (tested — 3 reasoning levels)
  • More large models coming

Cloud (open-source via API)

  • Open-source models via OpenRouter, Alibaba, and other affordable providers

💬 Want us to test a specific model? Open an issue or drop a comment — we prioritize community requests. New models are added as they release.

📄 License

MIT

AIgenteur/ClawEval

49

stars

0

commits

Python

primary language

Sep 10, 2026

updated

README

🦞 ClawEval — Can Your Local or Cloud LLM Actually Do the Job?

AIgenteur

The only deterministic benchmark that tests LLMs as real AI agents — not trivia, not chat, not vibes. 59 specialized roles. Hard, verifiable tasks. Every score is reproducible. Built for OpenClaw, Hermes Agent, and any model-agnostic agent framework.

🔥 New models added regularly. Star ⭐ this repo to get notified when new results drop.

📬 Subscribe to AIgenteur Dispatch — benchmark results, local AI guides, and agent workflows in your inbox.

🎓 Join AIgenteur Launchpad (free) — community for freelancers and entrepreneurs putting AI agents to work. Want a model tested? Post your request in the community.

🤖 Sister project: AIgenteur — open-source AI cofounder for solopreneurs. ClawEval grades the agents it dispatches.

💡 You'd be surprised how well some small, fast models handle sub-agent tasks — even on an RTX 3090 (we got ours for $799). Before you rent cloud GPUs, check the scores below.

Why ClawEval?

Most benchmarks tell you a model is "smart." ClawEval tells you if it can do the work — route tickets, review code, analyze financials, draft legal docs, plan sprints, and 54 more agent roles. Each test has an exact expected answer. No LLM-as-judge. No vibes.

Built for agent frameworks like OpenClaw and Hermes Agent — any system where you need to pick the right LLM backend for specialized sub-agent roles. If your framework is model-agnostic, ClawEval shows you exactly which model to plug in for each job.

⚡ NEW: Token Efficiency Leaderboard

🏆 Which models waste the fewest tokens? →

Two models both score 10/10 — but one used 500 tokens, the other 5,000. The efficient one saves 10× on compute, cost, and latency. We rank 35 models by tokens-per-point. Think modes use 10–20× more tokens for marginal gains. This metric is hardware-independent.

🚀 NEW: ClawEval v2 — 59 Agents, 1,220 Checkpoints, Zero Ceiling Effects

🏆 The definitive dense evaluation →

Phase F scored x/10 — everyone got 8+. ClawEval v2 upgrades all 59 agents to 15–30 checkpoints each with adversarial traps, sarcasm, and near-truths. Real separation, real rankings. First results (Kimi K2.6, GLM-5.1) dropping now.

🖥️ LOCAL Models — Run on YOUR Hardware

📄 OpenClaw Backend: Local on 3090 — The Complete Guide (PDF) — An interactive walkthrough of running the full OpenClaw backend locally on an RTX 3090. Covers setup, model selection, and performance tips with step-by-step calls to action.

We test quantized open-source models that fit on hardware you already own. Find out which model is best for each agent role before you commit your VRAM. We're also testing smaller models for max context and usability — even a 16GB GPU can run capable sub-agents.

🔴 8GB VRAM🟠 12GB VRAM🟡 16GB VRAM🟢 24GB VRAM🔵 64–96GB VRAM
Qwen3.5-0.8B Q4_K_MQwen3.5-9B Q4_K_MQwen3.5-9B Q4_K_MQwen3.6-35B-A3B Q4_K_MQwen3.5-122B-A10B NVFP4
Qwen3.5-2B Q4_K_MQwen3.5-35B-A3B Q4_K_MGPT-OSS-120B GGUF
Qwen3.5-4B Q4_K_M (250K ctx)Qwen3.5-27B Q4_K_M
✅ Tested✅ Testedllama.cppllama.cpp · SGLang · vLLMSGLang · vLLM · llama.cpp

📖 VRAM Guides: 8–16GB Small Models · 16GB · 24GB · 32GB · 48GB · 64GB · 96GB — Which models fit, context limits, speed estimates

🧩 Which Tested Models Fit on Your Hardware?

No discrete GPU? Unified-memory devices can run these models directly. Estimates are conservative — they account for OS overhead (~4–8 GB) and KV cache for reasonable context windows.

Model (as tested)WeightsMac Mini M4 (16 GB)Mac Mini M4 Pro (24 GB)Mac Mini M4 Pro (48 GB)Strix Halo (96 GB GPU)DGX Spark (96 GB GPU)
Qwen3.5-0.8B Q4_K_M~0.5 GB
Qwen3.5-2B Q4_K_M~1.5 GB
Gemma-4-E2B Q4_K_M~1.5 GB
Qwen3.5-4B Q4_K_M~2.8 GB
Gemma-4-E4B Q4_K_M~2.5 GB
Ministral-3B Q4_K_M~2 GB
Ministral-8B Q4_K_M~5 GB
Qwen3.5-9B Q4_K_M~6 GB
Ministral-14B Q4_K_M~9 GB⚠️ tight
Gemma-4-26B-A4B Q4_K_M~15 GB⚠️ tight
Qwen3.5-27B Q4_K_M~17 GB⚠️ tight
Gemma-4-31B Q4_K_M~19 GB⚠️ tight
Qwen3.5-35B-A3B NVFP4~20 GB
Qwen3.6-35B-A3B Q4_K_M~20 GB
GPT-OSS-120B GGUF~61 GB
Qwen3.5-122B-A10B NVFP4~70 GB
Nemotron-3-Super-120B-A12B NVFP4~70 GB
Mistral-Small-4-119B NVFP4~70 GB

⚠️ tight = model loads but leaves little room for KV cache / long context. Short prompts only.

Strix Halo = AMD Ryzen AI Max (128 GB LPDDR5X total, up to 96 GB allocatable to GPU). DGX Spark = NVIDIA GB10 (128 GB LPDDR5X unified, ~96 GB GPU-usable). Both can run 120B-class models with room for KV cache.

Mac Mini M4 Pro 64 GB and M4 Max 128 GB configs exist but are not listed — they slot between the columns above.

☁️ CLOUD Models — Run via API

Don't have a GPU? We also test open-source models hosted on cloud providers so you can compare local vs cloud performance at every agent role. Same tests, same scoring — find out if paying for cloud is worth it, or if your local setup already matches.

ProviderModelsStatus
Alibaba Coding PlanQwen3.5-Plus — 🥇 472/590 (80%) Phase F, 86/110 (78%) Phase G✅ Tested
Alibaba Coding PlanKimi K2.5 — 473/590 (80%) Phase F, 96/110 (87%) Phase G✅ Tested
Alibaba Coding PlanGLM-5 — 456/590 (77%) Phase F, 80/110 (73%) Phase G✅ Tested
Alibaba Coding PlanMiniMax-M2.5 — 455/590 (77%) Phase F, 78/110 (71%) Phase G✅ Tested
Ollama CloudMiniMax-M2.7 — 473/590 (80%) Phase F, 83/110 (75%) Phase G✅ Tested
OpenRouterStep-3.5-Flash — 438/590 (74%) Phase F, 66/110 (60%) Phase G✅ Tested
OpenRouterTrinity-Large — 444/590 (75%) Phase F, 73/110 (66%) Phase G✅ Tested
OpenRouterNemotron-3-Nano-30B — 453/590 (77%) Phase F, 70/110 (64%) Phase G✅ Tested
OpenRouterQwen3.6-Plus — 469/590 (79%) Phase F, 91/110 (83%) Phase G✅ Tested

💡 Kimi K2.5 leads Phase F at 473/590 (80%), closely followed by M2.7 and Qwen3.5-Plus. M2.7's reasoning tokens count against max_tokens, requiring 16-24k to avoid truncation.

📊 Detailed cloud comparisons:

⚡ Token Efficiency — Which Models Waste the Fewest Tokens?

Two models both score 10/10 — but one used 500 tokens, the other 5,000. The efficient one saves 10× on compute, cost, and latency. This is hardware-independent: it measures model behavior, not GPU speed.

RankModelVRAMScoreTokens/PointVerdict
🥇Trinity-Large☁️41439Ultra-lean
🥈Mistral-119B NT🔵 64GB41949Ultra-lean
🥉GLM-5 NT☁️42161Lean
4Qwen3.5-Plus NT☁️44263Lean
5Qwen3.5-35B-A3B NT🟢 24GB43276Efficient
6Kimi K2.5 Think☁️44380Efficient
...
32Qwen3.6-35B-A3B🟢 24GB430696Heavy
35Qwen3.5-35B-A3B Think🟢 24GB3951,524Extremely Heavy

💡 Key insight: Thinking models use 10–20× more tokens than NoThink variants for similar scores. For cost-sensitive agent pipelines, NoThink wins massively on efficiency.

📊 Full Token Efficiency Leaderboard (35 models) → — complete rankings, Think vs NoThink gaps, best per VRAM tier


🚀 ClawEval v2 — Full 59-Agent Dense Evaluation

The next generation of ClawEval testing. Every agent role now has a dense constraint test with 15–30 checkpoints — no more x/10 ceiling effects. This is the definitive leaderboard.

Phase F gave every model 8–10/10 on most roles. ClawEval v2 replaces that with granular percentage scoring across all 59 agents. Same roles, dramatically harder prompts, real separation between models.

MetricPhase F (v1)ClawEval v2
Tests59 roles × 10 pts59 roles × 15–30 checkpoints
Total checkpoints5901,220
Scoringx/10 (ceiling effect)Raw % (fine-grained)
Manual-only tests60 — fully automated
Discriminating powerLow (everyone scores 8+)High (real spread)

🏅 ClawEval v2 Leaderboard — Open-Weight Models

RankModelProviderScore%Perfect
🥇DeepSeek V4 Pro☁️ DeepSeek1060/122086.9%26
🥈DeepSeek V4 Flash☁️ DeepSeek1054/122086.4%23
🥉Kimi K2.5 Think☁️ Ollama1048/122085.9%24
4Kimi K2.7 Code☁️ Ollama1038/122085.1%24
5Laguna-M.1☁️ OpenRouter1034/122084.8%25
6Qwen3.5-Plus☁️ Alibaba1031/122084.5%26
7Qwen3.6-35B-A3B🖥️ Local1029/122084.3%24
8Kimi K2.6☁️ Ollama1028/122084.3%24
9Qwen3.5-122B-A10B☁️ Ollama1025/122084.0%21
10Gemma-4-31B☁️ Ollama1024/122083.9%25
11Cobuddy☁️ OpenRouter1023/122083.9%21
12Mistral-Large-3☁️ Ollama1021/122083.7%25
13GLM-5.1☁️ Ollama1020/122083.6%26
14Nemotron-3-Super Think☁️ Ollama1016/122083.3%20
15MiniMax-M2.7 Medium☁️ Ollama1014/122083.1%20
16Qwen3.6-27B🖥️ Local TQ41012/122083.0%26
17Nemotron-3-Super NoThink☁️ Ollama996/122081.6%21
18MiniMax-M3☁️ Ollama993/122081.4%25
19MiniMax-M2.7 Think☁️ Ollama993/122081.4%19
20Nemotron-3-Nano-Omni☁️ OpenRouter991/122081.2%20
21Gemma-4-E2B🖥️ Local981/122080.4%14
22GPT-OSS-120B☁️ Ollama979/122080.2%19
23Gemini-3.5-Flash☁️ Google (OpenRouter)978/122080.2%22
24Phi-4🖥️ Local Q8977/122080.1%17
25GLM-5.2☁️ Ollama957/122078.4%23
26Laguna-XS.2☁️ OpenRouter950/122077.9%20
27GLM-5 NoThink☁️ Ollama948/122077.7%25
28Nemotron-Nano-Omni🖥️ Local IQ4948/122077.7%20
29Kimi K2.5 NoThink☁️ Ollama935/122076.6%21
30Granite-4.1 30B🖥️ Local TQ4929/122076.1%15
31Granite-4.1 8B🖥️ Local Q4929/122076.1%14
32Gemma-4-31B🖥️ Local Q4927/122076.0%25
33GLM-5 Think☁️ Ollama927/122076.0%23
34Nemotron-Nano-Omni🖥️ Local Q4925/122075.8%19
35Trinity-Large-Think☁️ OpenRouter914/122074.9%21
36Nemotron-3-Nano-30B☁️ Ollama914/122074.9%19
37Ministral-3 8B☁️ Ollama906/122074.3%18
38Ministral-3 14B☁️ Ollama888/122072.8%17
39GPT-OSS-20B☁️ Ollama885/122072.5%19
40Ministral-3 8B🖥️ Local Q4884/122072.5%16
41Ministral-3 14B🖥️ Local Q4877/122071.9%18
42Gemma-4-E4B🖥️ Local867/122071.1%15
43Ministral-3 14B Think🖥️ Local Q4858/122070.3%19
44Granite-4.1 3B🖥️ Local Q4846/122069.3%12
45Ministral-3 3B☁️ Ollama844/122069.2%14
46Ministral-3 8B Think🖥️ Local Q4791/122064.8%10
47Ministral-3 3B🖥️ Local Q4760/122062.3%12
48RNJ-1-8B☁️ Ollama750/122061.5%18
49Ministral-3 3B Think🖥️ Local Q4704/122057.7%10
50Gemma-4-A4B🖥️ Local622/122051.0%10
51Qwen3.5-9B🖥️ Local543/122044.5%6
52Qwen3.5-4B🖥️ Local374/122030.7%4
53LFM2.5-350M🖥️ Local308/122025.2%2
54Qwen3.5-0.8B🖥️ Local58/12204.8%0
55Qwen3.5-2B🖥️ Local50/12204.1%0

📋 Gemini-3.5-Flash — Full 59-Agent Breakdown

Per-role performance across all 59 ClawEval v2 agents. 978/1,220 checkpoints (80.2%) — 22 roles at 100%, 12 roles below 50%.

TestAgent RoleScore%
H-01Router / Triage Agent29/3097%
H-02Input Validator / Sanitizer29/3097%
H-03Heartbeat / Health Monitor14/1593%
H-04Notification / Alert Agent21/3070%
H-05Sentiment Analysis Agent29/3097%
H-06FAQ Generation Agent15/15100%
H-07Translation Agent15/15100%
H-08Calendar / Scheduling Agent12/2060%
H-09Research / Web Search Agent30/30100%
H-10Content Writer / Blog Writer18/2090%
H-11Editor Agent29/3097%
H-12Content Planner / Strategist0/300%
H-13Email Drafting / Summarization43/4596%
H-14Document Summarization Agent15/15100%
H-15Meeting Notes / Transcription Agent35/35100%
H-16Social Media Scouting / Monitoring57/6095%
H-17Social Media Content Agent19/2095%
H-18News Aggregation Agent7/7100%
H-19Shopping / Price Comparison15/15100%
H-20Memory / Knowledge Management20/20100%
H-21RAG / Retrieval Agent4/1527%
H-22Data Analysis Agent14/1593%
H-23Website Scraping / Understanding15/15100%
H-24Image Description / Understanding20/20100%
H-25Customer Support Agent33/6055%
H-26Lead Scoring / Prospecting12/1580%
H-27Sprint / Project Summarizer15/15100%
H-28Transaction / Approval Agent19/2095%
H-29Home Automation Agent9/2045%
H-30Fitness / Health Tracking15/15100%
H-31Recipe / Cooking Agent15/15100%
H-32Personal Finance Tracking15/15100%
H-33SEO Optimization Agent4/1527%
H-34Landing Page Generator20/20100%
H-35Travel Planning Agent6/1540%
H-36Code Generation Agent30/30100%
H-37Code Review Agent15/15100%
H-38QA / Test Writing Agent13/1587%
H-39Task Planning / Decomposition1/186%
H-40Fact-Checking Agent28/3093%
H-41Critic / Review Agent20/20100%
H-42Market Research Agent14/1593%
H-43Synthesizer / Aggregator13/1587%
H-44Curriculum / Course Designer7/1547%
H-45Prototype Generator15/15100%
H-46DevOps Agent8/1553%
H-47Math / Logic Reasoning14/1593%
H-48STEM Research Analyst15/15100%
H-49Algorithm / Data Structure Explorer30/30100%
H-50Orchestrator / Manager Agent1/157%
H-51Software Architect Agent4/1527%
H-52Complex Debugger Agent13/1587%
H-53Legal Document Review6/1540%
H-54Medical / Health Analysis15/15100%
H-55Financial Analysis / Stock Research14/1593%
H-56Security Analyst Agent0/150%
H-57SRE / Incident Response13/1587%
H-58Book / Long-Form Writing19/2095%
H-59Compliance / Regulatory Agent2/1513%

📦 Evaluation Phases

ClawEval evaluates models across multiple phases of increasing difficulty:

PhaseFocusTestsScoring
A–CBasic role evaluation59 roles × system promptsQuality review
DHard prompts59 roles × adversarial promptsAutomated
EKiller tests12 precision tasksDeterministic (JSON, code exec, regex)
FRole-specific hard tests59 roles × deterministic promptsDeterministic (15+ scoring types)

Phase F — 59 Roles Across 5 Tiers

  • Tier 1 (8 roles): Router, Validator, Health Monitor, Notification, Sentiment, FAQ, Translation, Calendar
  • Tier 2 (27 roles): Research, Writer, Editor, Email, Summarization, Customer Support, Data Analysis, etc.
  • Tier 3 (11 roles): Code Gen, Code Review, QA Testing, Task Planning, Fact-Checking, Market Research, etc.
  • Tier 4 (3 roles): Math/Logic Reasoning, STEM Analysis, Algorithm Exploration
  • Tier 5 (10 roles): Orchestrator, Architect, Debugger, Legal, Medical, Financial, Security, SRE, etc.

Scoring Types

exact_json · json_numeric · keyword_detection · code_exec · constraint_check · error_count · action_items · news_dedup · recipe_scaling · compliance_issues · architecture_constraints · manual_review · and more

🚀 Quick Start

Prerequisites

python3 -m venv .venv
source .venv/bin/activate
pip install requests

Run Phase E (12 killer tests)

# Against llama.cpp server
python eval/run_phase_e.py \
  --base-url http://localhost:8080/v1 \
  --model my-model-name \
  --max-tokens 4000

# Against SGLang server (with thinking)
python eval/run_phase_e.py \
  --base-url http://192.168.1.2:8000/v1 \
  --model Qwen3.5-122B-think \
  --api-model /path/to/weights \
  --thinking-budget 16384 \
  --max-tokens 32000

# Disable thinking
python eval/run_phase_e.py \
  --base-url http://192.168.1.2:8000/v1 \
  --model Qwen3.5-122B-nothink \
  --api-model /path/to/weights \
  --nothink

Run Phase F (59 role tests)

python eval/run_phase_f.py \
  --base-url http://localhost:8080/v1 \
  --model my-model-name \
  --max-tokens 4000

# Run specific tests or tiers
python eval/run_phase_f.py ... --test-ids 1 2 3
python eval/run_phase_f.py ... --tier 3

Results

Results are saved to eval/test_results/<model-name>/phase_e/ and phase_f/ directories, including:

  • Individual response text files per test
  • phase_e_scores.json / phase_f_scores.json with full scoring breakdown

🔧 Key Features

  • OpenAI-compatible API: Works with any server exposing /v1/chat/completions (llama.cpp, SGLang, vLLM, etc.)
  • SGLang thinking control: --thinking-budget and --nothink flags for reasoning token management
  • Think-tag stripping: Automatically strips <think> tags from responses for clean scoring
  • Flexible JSON scoring: Case-insensitive keys, nested dict traversal, numeric tolerance
  • Deterministic: Every test has an exact expected answer — no LLM-as-judge

📁 Repo Structure

ClawEval/
├── eval/
│   ├── run_phase_e.py          # Phase E runner (12 killer tests)
│   ├── run_phase_f.py          # Phase F runner (59 role tests)
│   ├── phase_e_prompts.py      # Phase E test definitions
│   ├── phase_f/                # Phase F test definitions (5 tiers)
│   │   ├── __init__.py
│   │   ├── tier1.py ... tier5.py
│   ├── role_prompts.py         # 59 role system prompts
│   ├── hard_prompts.py         # Phase D adversarial prompts
│   └── test_results/           # All model evaluation results
├── docs/                       # VRAM tier guides & model selection
│   ├── 16GB, 24GB, 32GB, 48GB, 64GB, 96GB tier guides
│   └── Subagent type reference
├── RESULTS.md                  # Detailed per-role score comparison
└── README.md

🗺️ Model Testing Roadmap

This is a living benchmark. We're continuously adding new models as they release. Here's what's on the radar — and we take requests.

16GB VRAM tier (coming soon)

  • Qwen3-8B, Qwen3.5-14B
  • Phi-4 14B, Phi-4-mini 3.8B
  • Mistral Small 3.1 24B (tight fit)
  • GPT-OSS-20B MXFP4
  • Gemma 3 12B
  • Ministral 3B (ultra-fast routing/triage)

24GB VRAM tier

  • ✅ Qwen3.5-35B-A3B Q4_K_M (tested)
  • ✅ Qwen3.5-27B Q4_K_M (tested)
  • Gemma 3 27B QAT
  • More MoE models as they release

64–96GB VRAM tier

  • ✅ Qwen3.5-122B-A10B NVFP4 (tested)
  • ✅ GPT-OSS-120B GGUF · llama.cpp (tested — 3 reasoning levels)
  • More large models coming

Cloud (open-source via API)

  • Open-source models via OpenRouter, Alibaba, and other affordable providers

💬 Want us to test a specific model? Open an issue or drop a comment — we prioritize community requests. New models are added as they release.

📄 License

MIT

Languages

Python

100.0%