Hanno-Labs/decision-bench-results

Python

1

76 commits

updated Sep 28, 2026

See the code

README


benchmark: decision-bench type: evaluation submission_name: DecisionBench

DecisionBench Results

This repository is the reviewed, append-only result registry for DecisionBench.

Reference
πŸ“ˆ LeaderboardBrowse reviewed model results
πŸ“š DecisionBenchRun evaluations and load result records
🧾 Submission guideValidate and submit a new result
πŸ› IssuesReport bugs or request features for any DecisionBench component

Layout

results/<model-name>/<immutable-revision>/
β”œβ”€β”€ model_meta.json
β”œβ”€β”€ DecisionBench.json
└── DecisionBench--compact.json  # optional tagged run

Each compact record contains the reviewed metrics and immutable model and dataset identities. A submission may also point to complete row-level artifacts in durable storage. Unsupported and error rows count as misses in the primary leaderboard score.

Result tags are optional. Existing results have no tags. A result tagged compact uses a shorter rendering of the same benchmark rows, preserving candidate meaning, gold labels, and scoring. Tagged results appear on the same leaderboard with a visible tag and can coexist with the untagged result for the same model revision.

Each model declares its own reviewed model_type: decision-model for checkpoints trained across the benchmark's noul, choice, and score primitives with a native decision output; language-model for text-generating models; or classifier for fixed-purpose class, relevance, or scalar scorers. Architecture names and API access do not determine this field. Stock language models exposed through Jev-compatible servers remain language-model; the server contract does not turn their weights into a trained decision checkpoint. The leaderboard can show underlying weights when the result URL identifies a base checkpoint.

Model names use an owner/name identity. For serving recipes without their own Hugging Face checkpoint, owner credits the upstream recipe project. The url and revision identify the base weight snapshot when one is recorded; older recipe submissions may instead identify their source-code revision. The table below names the evaluated base checkpoint where verified. A recipe's name must not imply that its author published separate model weights.

For quantized variants, parameter_count is the logical count of the source model, not the number of packed storage tensors in the quantized checkpoint.

Recipe resultUpstream projectEvaluated base checkpoint
ekzhang/openjev-sglangekzhang/openjev-sglangnvidia/Qwen3.6-35B-A3B-NVFP4
razorback16/openjevrazorback16/openjevnvidia/diffusiongemma-26B-A4B-it-NVFP4
TheoLeeCJ/SemIfTheoLeeCJ/SemIfQwen/Qwen3.5-4B
Octalab-Inc/jqvOctalab-Inc/jqvQwen/Qwen3-32B
zhengxuyu/litjevzhengxuyu/litjevQwen/Qwen3.8-27B
kshetrajna12/reflex-27bkshetrajna12/reflexQwen/Qwen3.8-27B
featherless-ai/simplejev-qwen3.6-35b-a3bfeatherless-ai/simple-jevQwen/Qwen3.6-35B-A3B
featherless-ai/simplejev-qwen3.8-27bfeatherless-ai/simple-jevQwen/Qwen3.8-27B

Hanno-Labs/decision-bench-results

Python

1

76 commits

updated Sep 28, 2026

See the code

README


benchmark: decision-bench type: evaluation submission_name: DecisionBench

DecisionBench Results

This repository is the reviewed, append-only result registry for DecisionBench.

Reference
πŸ“ˆ LeaderboardBrowse reviewed model results
πŸ“š DecisionBenchRun evaluations and load result records
🧾 Submission guideValidate and submit a new result
πŸ› IssuesReport bugs or request features for any DecisionBench component

Layout

results/<model-name>/<immutable-revision>/
β”œβ”€β”€ model_meta.json
β”œβ”€β”€ DecisionBench.json
└── DecisionBench--compact.json  # optional tagged run

Each compact record contains the reviewed metrics and immutable model and dataset identities. A submission may also point to complete row-level artifacts in durable storage. Unsupported and error rows count as misses in the primary leaderboard score.

Result tags are optional. Existing results have no tags. A result tagged compact uses a shorter rendering of the same benchmark rows, preserving candidate meaning, gold labels, and scoring. Tagged results appear on the same leaderboard with a visible tag and can coexist with the untagged result for the same model revision.

Each model declares its own reviewed model_type: decision-model for checkpoints trained across the benchmark's noul, choice, and score primitives with a native decision output; language-model for text-generating models; or classifier for fixed-purpose class, relevance, or scalar scorers. Architecture names and API access do not determine this field. Stock language models exposed through Jev-compatible servers remain language-model; the server contract does not turn their weights into a trained decision checkpoint. The leaderboard can show underlying weights when the result URL identifies a base checkpoint.

Model names use an owner/name identity. For serving recipes without their own Hugging Face checkpoint, owner credits the upstream recipe project. The url and revision identify the base weight snapshot when one is recorded; older recipe submissions may instead identify their source-code revision. The table below names the evaluated base checkpoint where verified. A recipe's name must not imply that its author published separate model weights.

For quantized variants, parameter_count is the logical count of the source model, not the number of packed storage tensors in the quantized checkpoint.

Recipe resultUpstream projectEvaluated base checkpoint
ekzhang/openjev-sglangekzhang/openjev-sglangnvidia/Qwen3.6-35B-A3B-NVFP4
razorback16/openjevrazorback16/openjevnvidia/diffusiongemma-26B-A4B-it-NVFP4
TheoLeeCJ/SemIfTheoLeeCJ/SemIfQwen/Qwen3.5-4B
Octalab-Inc/jqvOctalab-Inc/jqvQwen/Qwen3-32B
zhengxuyu/litjevzhengxuyu/litjevQwen/Qwen3.8-27B
kshetrajna12/reflex-27bkshetrajna12/reflexQwen/Qwen3.8-27B
featherless-ai/simplejev-qwen3.6-35b-a3bfeatherless-ai/simple-jevQwen/Qwen3.6-35B-A3B
featherless-ai/simplejev-qwen3.8-27bfeatherless-ai/simple-jevQwen/Qwen3.8-27B

Languages

Python

99.6%