multimodalart/jev-decision-index

Space

Jev Reproductions Tracker

96

60 commits

updated Sep 22, 2026

See the code

README

Jev Reproductions Tracker

Who is rebuilding TypeSafe's Jev (System One / RLCD) in the open?

This static Space opens on the Decision Index leaderboard; the News tab tracks the artifacts in one combined grid, color-coded by kind:

  • Decoding: parallel constrained decoding on stock models (inference technique, no new weights)
  • Diffusion: text diffusion models run in a "Jev mode"
  • Trained: Jev-like scoring heads and fine-tunes, weights often on the Hub, promised models listed last
  • Prior art: "this already exists" claims
  • Explainers: architecture speculation, explainers, benchmarks and roundups

Cards sort by a trending score: ♥ likes on X + 5 × GitHub stars + 8 × Hub likes + views ÷ 500, plus a recency credit (up to 8000, halving daily) so a fresh release is not buried under older, louder ones. The credit is scaled by the item's own engagement, so a brand-new entry nobody has reacted to yet cannot climb on its date alone. Cards added within a day of the metrics snapshot get a green new badge. Use the category and "has" chips to filter, or switch the sort.

…plus a section on what is still not in the open: TypeSafe's weights, the RLCD algorithm, and any open model matching Jev's calibration claims.

Hugging Face models and Spaces are first-class artifacts here, not only GitHub repos.

Decision Index 0.1 (index.html)

The Index tab (switch at the top of every page) is the leaderboard: every open reproduction that finished the frozen 132,422-request suite, scored with one number, the Decision Index, plus per-category radars against Jev and a full page per model (index.html?model=<engine>; ?model=jev is Jev's own report).

All numbers come from one static bundle, data/index.json, written by evaluation/reproductions/build_leaderboard.py in the typesafe-diffusion-lab checkout from capability-indices.json, entrant-metadata.json, each run's benchmark-summary.json and the Jev release report. Re-run that script and commit the JSON to refresh the page; the HTML never needs to change for a data update.

the answered-rate explanations come from evaluation/reproductions/answer_gaps.py, which mines every run's refusal messages into one line per cause and writes answer-gaps.json for the build to pick up.

Editorial choices baked into the build: the six interactive environments in the frozen panel are left out of every model's index (Jev included) because no reproduction has run them, so the index is a 19-benchmark panel across five areas (ChessBench folded into Knowledge & Reasoning, Language split into Language Understanding and Retrieval & Classification); iSarcasmEval's headline is track A (English sarcasm F1) with the other tracks listed under each result; only the balanced raw index is shown, the chance-normalized and breadth variants are computed but hidden.

Regenerating the social image

og-index.html is the 1200×630 stage that produces og.png, the Space thumbnail (title lockup plus the Decision Index bar chart, read live from data/index.json; render at 2× with headless Chrome after regenerating the data). og-news.html is the older stage for the News tab and produces og-news.png.

Contributing

All news data lives in one JS array (ITEMS) near the bottom of news.html (the News tab; the Index tab is index.html). Open a PR on this Space to add or correct an entry. Please include the announcement post on X if there is one, and the Hub / GitHub link.

Engagement metrics are a snapshot (2026-09-20) from public post metadata and will drift.

Credits

Visual identity and outlined Huggies from HF Huggiverse (Chunte/HFBA). Unofficial and community-maintained; not affiliated with TypeSafe AI.

static

Contributors

multimodalart

20 commits

apolinario

19 commits

CC
Claude Code

12 commits

fredreick

3 commits

multimodalart/jev-decision-index

Space

Jev Reproductions Tracker

96

60 commits

updated Sep 22, 2026

See the code

README

Jev Reproductions Tracker

Who is rebuilding TypeSafe's Jev (System One / RLCD) in the open?

This static Space opens on the Decision Index leaderboard; the News tab tracks the artifacts in one combined grid, color-coded by kind:

  • Decoding: parallel constrained decoding on stock models (inference technique, no new weights)
  • Diffusion: text diffusion models run in a "Jev mode"
  • Trained: Jev-like scoring heads and fine-tunes, weights often on the Hub, promised models listed last
  • Prior art: "this already exists" claims
  • Explainers: architecture speculation, explainers, benchmarks and roundups

Cards sort by a trending score: ♥ likes on X + 5 × GitHub stars + 8 × Hub likes + views ÷ 500, plus a recency credit (up to 8000, halving daily) so a fresh release is not buried under older, louder ones. The credit is scaled by the item's own engagement, so a brand-new entry nobody has reacted to yet cannot climb on its date alone. Cards added within a day of the metrics snapshot get a green new badge. Use the category and "has" chips to filter, or switch the sort.

…plus a section on what is still not in the open: TypeSafe's weights, the RLCD algorithm, and any open model matching Jev's calibration claims.

Hugging Face models and Spaces are first-class artifacts here, not only GitHub repos.

Decision Index 0.1 (index.html)

The Index tab (switch at the top of every page) is the leaderboard: every open reproduction that finished the frozen 132,422-request suite, scored with one number, the Decision Index, plus per-category radars against Jev and a full page per model (index.html?model=<engine>; ?model=jev is Jev's own report).

All numbers come from one static bundle, data/index.json, written by evaluation/reproductions/build_leaderboard.py in the typesafe-diffusion-lab checkout from capability-indices.json, entrant-metadata.json, each run's benchmark-summary.json and the Jev release report. Re-run that script and commit the JSON to refresh the page; the HTML never needs to change for a data update.

the answered-rate explanations come from evaluation/reproductions/answer_gaps.py, which mines every run's refusal messages into one line per cause and writes answer-gaps.json for the build to pick up.

Editorial choices baked into the build: the six interactive environments in the frozen panel are left out of every model's index (Jev included) because no reproduction has run them, so the index is a 19-benchmark panel across five areas (ChessBench folded into Knowledge & Reasoning, Language split into Language Understanding and Retrieval & Classification); iSarcasmEval's headline is track A (English sarcasm F1) with the other tracks listed under each result; only the balanced raw index is shown, the chance-normalized and breadth variants are computed but hidden.

Regenerating the social image

og-index.html is the 1200×630 stage that produces og.png, the Space thumbnail (title lockup plus the Decision Index bar chart, read live from data/index.json; render at 2× with headless Chrome after regenerating the data). og-news.html is the older stage for the News tab and produces og-news.png.

Contributing

All news data lives in one JS array (ITEMS) near the bottom of news.html (the News tab; the Index tab is index.html). Open a PR on this Space to add or correct an entry. Please include the announcement post on X if there is one, and the Hub / GitHub link.

Engagement metrics are a snapshot (2026-09-20) from public post metadata and will drift.

Credits

Visual identity and outlined Huggies from HF Huggiverse (Chunte/HFBA). Unofficial and community-maintained; not affiliated with TypeSafe AI.

static

Contributors

multimodalart

20 commits

apolinario

19 commits

CC
Claude Code

12 commits

fredreick

3 commits