This repository is the reviewed, append-only result registry for DecisionBench.
| Reference | |
|---|---|
| π Leaderboard | Browse reviewed model results |
| π DecisionBench | Run evaluations and load result records |
| π§Ύ Submission guide | Validate and submit a new result |
| π Issues | Report bugs or request features for any DecisionBench component |
results/<model-name>/<immutable-revision>/
βββ model_meta.json
βββ DecisionBench.json
βββ DecisionBench--compact.json # optional tagged run
Each compact record contains the reviewed metrics and immutable model and dataset identities. A submission may also point to complete row-level artifacts in durable storage. Unsupported and error rows count as misses in the primary leaderboard score.
Result tags are optional. Existing results have no tags. A result tagged compact uses a
shorter rendering of the same benchmark rows, preserving candidate meaning, gold labels,
and scoring. Tagged results
appear on the same leaderboard with a visible tag and can coexist with the untagged
result for the same model revision.
Each model declares its own reviewed model_type: decision-model for checkpoints trained across
the benchmark's noul, choice, and score primitives with a native decision output;
language-model for text-generating models; or classifier for fixed-purpose class, relevance, or
scalar scorers. Architecture names and API access do not determine this field.
Stock language models exposed through Jev-compatible servers remain language-model; the
server contract does not turn their weights into a trained decision checkpoint. The leaderboard
can show underlying weights when the result URL identifies a base checkpoint.
Model names use an owner/name identity. For serving recipes without their own Hugging Face
checkpoint, owner credits the upstream recipe project. The url and revision identify the
base weight snapshot when one is recorded; older recipe submissions may instead identify their
source-code revision. The table below names the evaluated base checkpoint where verified. A
recipe's name must not imply that its author published separate model weights.
For quantized variants, parameter_count is the logical count of the source model, not
the number of packed storage tensors in the quantized checkpoint.
| Recipe result | Upstream project | Evaluated base checkpoint |
|---|---|---|
ekzhang/openjev-sglang | ekzhang/openjev-sglang | nvidia/Qwen3.6-35B-A3B-NVFP4 |
razorback16/openjev | razorback16/openjev | nvidia/diffusiongemma-26B-A4B-it-NVFP4 |
TheoLeeCJ/SemIf | TheoLeeCJ/SemIf | Qwen/Qwen3.5-4B |
Octalab-Inc/jqv | Octalab-Inc/jqv | Qwen/Qwen3-32B |
zhengxuyu/litjev | zhengxuyu/litjev | Qwen/Qwen3.8-27B |
kshetrajna12/reflex-27b | kshetrajna12/reflex | Qwen/Qwen3.8-27B |
featherless-ai/simplejev-qwen3.6-35b-a3b | featherless-ai/simple-jev | Qwen/Qwen3.6-35B-A3B |
featherless-ai/simplejev-qwen3.8-27b | featherless-ai/simple-jev | Qwen/Qwen3.8-27B |
Python
99.6%
This repository is the reviewed, append-only result registry for DecisionBench.
| Reference | |
|---|---|
| π Leaderboard | Browse reviewed model results |
| π DecisionBench | Run evaluations and load result records |
| π§Ύ Submission guide | Validate and submit a new result |
| π Issues | Report bugs or request features for any DecisionBench component |
results/<model-name>/<immutable-revision>/
βββ model_meta.json
βββ DecisionBench.json
βββ DecisionBench--compact.json # optional tagged run
Each compact record contains the reviewed metrics and immutable model and dataset identities. A submission may also point to complete row-level artifacts in durable storage. Unsupported and error rows count as misses in the primary leaderboard score.
Result tags are optional. Existing results have no tags. A result tagged compact uses a
shorter rendering of the same benchmark rows, preserving candidate meaning, gold labels,
and scoring. Tagged results
appear on the same leaderboard with a visible tag and can coexist with the untagged
result for the same model revision.
Each model declares its own reviewed model_type: decision-model for checkpoints trained across
the benchmark's noul, choice, and score primitives with a native decision output;
language-model for text-generating models; or classifier for fixed-purpose class, relevance, or
scalar scorers. Architecture names and API access do not determine this field.
Stock language models exposed through Jev-compatible servers remain language-model; the
server contract does not turn their weights into a trained decision checkpoint. The leaderboard
can show underlying weights when the result URL identifies a base checkpoint.
Model names use an owner/name identity. For serving recipes without their own Hugging Face
checkpoint, owner credits the upstream recipe project. The url and revision identify the
base weight snapshot when one is recorded; older recipe submissions may instead identify their
source-code revision. The table below names the evaluated base checkpoint where verified. A
recipe's name must not imply that its author published separate model weights.
For quantized variants, parameter_count is the logical count of the source model, not
the number of packed storage tensors in the quantized checkpoint.
| Recipe result | Upstream project | Evaluated base checkpoint |
|---|---|---|
ekzhang/openjev-sglang | ekzhang/openjev-sglang | nvidia/Qwen3.6-35B-A3B-NVFP4 |
razorback16/openjev | razorback16/openjev | nvidia/diffusiongemma-26B-A4B-it-NVFP4 |
TheoLeeCJ/SemIf | TheoLeeCJ/SemIf | Qwen/Qwen3.5-4B |
Octalab-Inc/jqv | Octalab-Inc/jqv | Qwen/Qwen3-32B |
zhengxuyu/litjev | zhengxuyu/litjev | Qwen/Qwen3.8-27B |
kshetrajna12/reflex-27b | kshetrajna12/reflex | Qwen/Qwen3.8-27B |
featherless-ai/simplejev-qwen3.6-35b-a3b | featherless-ai/simple-jev | Qwen/Qwen3.6-35B-A3B |
featherless-ai/simplejev-qwen3.8-27b | featherless-ai/simple-jev | Qwen/Qwen3.8-27B |
Python
99.6%