stefanwebb/meta-awesome-jev

The Awesome List of Awesome Jev Lists

Python

2

4 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Meta Analysis of "Awesome Jev" GitHub Repos (r/LocalLLaMA)

I noticed that there are heaps of Awesome Jev resource list posts popping up on GitHub (e.g. https://github.com/yibie/awesome-jev) so I thought why not ask Claude to pull them all in and do a meta analysis of insights and applications. Sharing in case it's of interest:…

0

Oct 1, 2026

README

Meta Awesome Jev-like Models Awesome

A guide to the community of Jev-like models. These are "System One" decision models: you send them a state and some typed questions, and they return calibrated choices, scores and yes/no probabilities instead of generated text.

Jev introduced the /v1/systemone interface in September 2026. Within weeks it was served by a growing family of hosted and open models, including Mercury Decide, Liquid d1, Solar Decide, Laya, Kev, Decider and more.

This list merges, de-duplicates, cross-checks and summarizes 106 community "awesome-jev" repositories.

18,330 unique links from 106 lists · 5,410 cited by 3 or more lists · 2,533 GitHub repos verified live · sources last pulled 2026-09-30 (sources.md)

★ = GitHub stars when verified on the pull date · 📚N = the number of the 106 source lists that cite the entry. 📚 is the cross-list consensus signal: an entry that many independently curated lists include has been vetted many times.


Contents

Deep dives in this repo:

docs/open-models.mdThe Jev-like model family: open replications, local runtimes and hosted alternatives, with caveats
docs/patterns.mdQuestion design, composition patterns, thresholds, economics, security, a production checklist and anti-patterns
docs/evidence.mdEvery independent benchmark, calibration audit, robustness probe and negative result, with numbers
docs/what-is-jev.mdA reference for Jev, the original model: specs, API, timeline, and contradictions between sources, resolved
docs/papers.md~35 papers on Jev and Jev-like models, plus the research lineage
docs/source-lists.mdA review of all 106 source lists: which to read for what, and which to treat with caution
catalog/The complete union of every link from every list, in 21 categories, ranked by consensus

Jev-like models in 60 seconds

A Jev-like, or System One, model is not a chat model, and it never writes text. You send it two things in one request:

  • a state: the evidence (a string, a JSON object, or an array of text);
  • a map of typed questions about that state.

It answers every question in parallel, in a single call. Each answer is a probability distribution that your code can threshold.

Every System One model in the family speaks the same /v1/systemone wire format. Here is a support ticket with one question of each type:

POST /v1/systemone
{
  "model": "<model id>",
  "state": {
    "ticket": "I was charged twice for my order last week and nobody has answered my emails. I want my money back today."
  },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this ticket?",
      "criteria": {
        "billing":   "Charges, refunds, invoices",
        "technical": "Bugs, outages, errors",
        "other":     "Anything else"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Annoyed but civil", "Angry or threatening to leave"]
    },
    "wants_refund": {
      "type": "noul",
      "instructions": "Is the customer asking for a refund?"
    }
  }
}

The response contains one typed answer per question:

{
  "model": "<resolved model version>",
  "answers": {
    "team":         { "type": "choice", "choice": "billing",
                      "probabilities": { "billing": 0.95, "technical": 0.02, "other": 0.03 },
                      "confidence": 0.93 },
    "frustration":  { "type": "score", "score": 1.62,
                      "legend": { "0": "Calm", "1": "Annoyed but civil", "2": "Angry or threatening to leave" },
                      "probabilities": { "0": 0.03, "1": 0.32, "2": 0.65 }, "confidence": 0.71 },
    "wants_refund": { "type": "noul", "noul": 0.97 }
  },
  "usage": { "input_tokens": 214, "output_tokens": 5 }
}

Your code then decides what to do with the answers:

a = response["answers"]
if a["team"]["confidence"] >= 0.8:          # "the answer is what; confidence is whether to act"
    route_to(a["team"]["choice"])
else:
    send_to_human_triage()
if a["wants_refund"]["noul"] > 0.9 and a["frustration"]["score"] >= 1.5:
    escalate_priority()
Question typeAsksYou supplyYou get back
Choice"Which one?"an instructions string and criteria as a map of option → description (up to 255 options; always include an other)choice, probabilities per option, confidence
Score"Where on this ordered scale?"an instructions string and criteria as an ordered list of 2–10 levelsscore (probability-weighted and 0-indexed, so it can fall between levels), legend, probabilities, confidence
Noul"Is this true?"an instructions stringnoul: the probability, from 0 to 1, that the answer is yes. There is no confidence field, and 0.5 means "can't tell", not "medium".

→ The reference for the original model and the question-design guide go deeper.

Quick Start: Jev-like inference for free

This is how to use a Jev-like (System One) model for free. Mercury Decide is Inception's System One decision model, served on OpenRouter as inception/mercury-decide:free. Details from its OpenRouter page:

  • $0 input and $0 output. Free endpoints are rate-limited.
  • 32,768-token context.
  • About 0.42 s p50 latency and up to ~14 decisions per second.
  • The same /v1/systemone question schema shown above.

1. Get a free OpenRouter API key at openrouter.ai/settings/keys.

2. Call the Decisions API. Mercury Decide is a decisions model, so it uses OpenRouter's Decisions API, not /chat/completions. OpenAI-style chat SDKs won't work with it.

export OPENROUTER_API_KEY=sk-or-...

curl https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "inception/mercury-decide:free",
    "state": "I was charged twice for my order and want my money back.",
    "questions": {
      "team":   { "type": "choice", "instructions": "Which team should handle this ticket?",
                  "criteria": { "billing": "Charges and refunds", "technical": "Bugs and outages", "other": "Anything else" } },
      "refund": { "type": "noul",   "instructions": "Is the customer asking for a refund?" }
    }
  }'

3. Or call it from Python (pip install requests; no other SDK is needed):

import os, requests

resp = requests.post(
    "https://openrouter.ai/api/alpha/decisions",
    headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
    json={
        "model": "inception/mercury-decide:free",
        "state": {"ticket": "I was charged twice for my order and want my money back."},
        "questions": {
            "team": {"type": "choice", "instructions": "Which team should handle this ticket?",
                     "criteria": {"billing": "Charges and refunds", "technical": "Bugs and outages",
                                  "other": "Anything else"}},
            "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
            "urgency": {"type": "score", "instructions": "How urgent is this ticket?",
                        "criteria": ["Can wait a week", "Handle within a day", "Handle within the hour"]},
        },
    },
    timeout=30,
)
resp.raise_for_status()
answers = resp.json()["answers"]
print(answers["team"]["choice"], answers["team"]["confidence"], answers["refund"]["noul"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"])  # e.g. 1.05 {'0': 0.04, '1': 0.87, '2': 0.09}

Tips:

  • Already have code for another /v1/systemone client? OpenRouter also accepts the same request body at https://openrouter.ai/api/v1/systemone, so you can point an existing client's base URL at https://openrouter.ai/api.
  • Batch questions. Put every independent question about a state into one request. It costs the same as a single question, so use the batching to stay inside the free rate limit.
  • Graduate when you need more. Move to a paid decision model, or self-host an open one (Laya, Kev or Ollaya), when you need higher limits or local data. Most of these keep the same schema; see the model family.
  • Evaluate on your own labelled examples before trusting thresholds. Calibration doesn't transfer between models, so a 0.8 cut-off tuned on one model means nothing on another.

What 106 lists taught us: key insights

The same lessons show up again and again across the lists. Most measurements were taken on Jev, the first and most-tested model, and are labelled as such. The lessons about how to build apply to the whole family. Numbers link to the underlying studies in docs/evidence.md.

  1. Jev-like models are a System One decision layer, not a model swap. The winning architecture is always the same: code builds the candidates, the decision model picks or scores, code acts. An LLM is called only for text or open reasoning. Examples:
    • Browser agents let the model choose the element and use an LLM only to type.
    • Games give the model the legal moves.
    • Extraction lets regex find spans and the model select one.
  2. The 100× speedups are real only against chains of LLM calls. Jev's launch figures ("193.6× faster / 444.6× cheaper") are a best case against the slowest comparator; the average across eight setups is 97.8× / 149.2×. A single question is only ~1.5–3× faster than a small chat model. The order-of-magnitude wins come from batching many questions per call, where latency is flat up to ~25 questions and batching 13 questions was 12.2× cheaper, and from deleting LLM calls that never needed an LLM.
  3. Question design is the whole game.
    • Splitting one question into atomic Nouls took a task from 48% → 98%, and phishing detection from 62.6% → 95.0%.
    • Rewriting the criteria moved accuracy from 70% → 96%.
    • Renaming options from 0/1 to no/yes cut AUC from .81 to .58.
    • Always include an "other/unknown" option. Removing it took one benchmark from 0.95 to 0.00.
  4. Keep computation in code. Counting (33–65%), arithmetic, dates, multi-hop reasoning and open-ended generation are documented weak spots. Jev's own jaggedness page lists them, and independent tests on open replicas agree. Use one Noul per item and sum in code.
  5. Confidence is a routing signal, not a guarantee.
    • The top-confidence band is very reliable (48/48 correct at ≥ 0.9 in one audit).
    • "Confidently wrong" still happens, and Choice and Score skew overconfident.
    • Calibration does not transfer across datasets, languages, phrasings, model versions or models.
    • Fit per-question thresholds on your own labels, and pin a model version.
  6. Cascades are the most consistent cost/accuracy win, with one catch. Accepting confident decisions and escalating the rest to an LLM matched frontier accuracy at ~¼ the cost in several studies. But LLMs repeated ~96% of Jev's confident errors, so cascades mostly save money rather than add accuracy. Price the fallback: one production replay was 96% cheaper per call but 4% more expensive overall.
  7. It is not a security boundary. Typed output blocks format attacks, not decision manipulation. Blunt injected commands mostly fail. Evidence-shaped text flips decisions: fake approvals, editor's notes and fluent context flipped 61% of correct answers in JevOut, and 65–73% on open clones. Put deterministic rules first, keep humans on irreversible actions, and track taint.
  8. Audit what tools send. Hands-on audits found community tools leaking secrets, sending .pem files and screenshots, exposing keys, and failing open on API errors. Check data egress and fail-closed behaviour before installing.
  9. Coding agents are the dominant use case. The largest clusters across all lists are context compaction, model/effort routing, tool-call gates, "done" verification, and skill/rule pruning. Compaction is the most contested of these: several careful evaluations (Hermes Agent, jev-use) did not adopt it.
  10. Classical baselines are still strong. Trained small classifiers (bge-small + LR at 93% on Banking77, TF-IDF on spam) and regex often match or beat decision models. The durable edges of System One models are zero training, cost, latency, and robustness under distribution drift: on drifted spam, Jev held at 97.3% while TF-IDF fell to 72.5%.
  11. The family grew fast, but "compatible" ≠ "equivalent". Within days, dozens of open models and /v1/systemone servers appeared (Laya, Kev, SemIf, NanoJev, Decider, Ollaya), followed by hosted alternatives (Mercury Decide, Liquid d1, Solar Decide). The shared schema makes them drop-ins for your code, not for each other's accuracy or calibration. Most "beats Jev" claims are in-distribution, so measure on your own data.
  12. Most of the evidence is young. Nearly all numbers are author-reported, small-n, and from the family's first weeks, and the ~29 arXiv papers are unreplicated preprints. The best lists label every number (vendor / author-reported / independent), and this one does too.

Reference docs and SDKs

Jev's documentation is the most complete public description of the /v1/systemone interface: its question types, confidence semantics, composition patterns and cookbooks. Because Jev-like models share the schema, most of it applies to every System One model.

Interface docs

Reference SDKs and tools

  • Python: typesafe-sdk-python ★257 · 📚59. JS/TS: typesafe-sdk-js ★260 · 📚57. Both are /v1/systemone clients. Set the base-URL environment variable to point either one at any compatible server (local Kev, Decider or Ollaya; OpenRouter; and others).
  • system-one-adapter-python ★368 · 📚61 — The same client interface answered by an ordinary LLM. Use it for A/B baselines and fallback.
  • skills ★2,491 · 📚68 — An agent skill that teaches coding agents how to pick a question type and design questions.
  • For model-agnostic clients across many languages, see Community SDKs.

Patterns and cookbooks (the full annotated index covers all ~18)


Ecosystem

Hand-picked from the consensus of the source lists: mostly entries cited by many lists, plus a few high-signal ones that fewer lists noticed. Each section links to its full catalog page.

Jev-like models: hosted, open and local

The System One model family itself. Many members serve /v1/systemone, so the same client code works across them. docs/open-models.md maps the landscape and its caveats.

Hosted

  • Mercury Decide (Inception) — A structured decision model on OpenRouter, free, with a 32K context and the /v1/systemone schema. See the Quick Start.
  • Jev (TypeSafe) — The original and most-tested model, with a reference.
  • Liquid AI d1 — Serves /decisions/v1/systemone and has a free tier. Reportedly #1 on the Jev Decision Index.
  • Upstage Solar Decide — Same schema, 512K context.
  • Respan Span-01 — A System One model for behaviour monitoring.
  • The OpenAI Decisions API (Luna, preview).

Open and local

Framework and platform integrations

System One decision models landed inside mainstream frameworks within days of Jev's launch, usually as a new evaluate / decide / classify model type alongside generate. Most of these integrations target the shared /v1/systemone shape.

Community SDKs and clients

Every major language had a community SDK within a week. Check the last commit date before you depend on one.

MCP servers

Agent skills and plugins

Coding-agent tooling

The largest cluster in the ecosystem. It covers Claude Code, Codex, Pi, Hermes, OpenCode and Cursor.

Context compaction and pruning. This is contested, so read the evidence first.

Model and effort routing. Route at session boundaries and run in shadow mode first.

Guardrails and tool-call gates. Put deterministic rules first, and choose fail-open or fail-closed explicitly.

Done-verification, review and semantic lint

Browser, desktop and mobile

The shared recipe: build a numbered list of what's on the page or screen (accessibility tree, DOM or OCR), let the decision model pick the operation and target, let code execute, and call an LLM only when text must be typed. Most decision models are text-only and can't see pixels.

Search, RAG and data

Evaluation, calibration and QA tooling

  • abhixhek/jevcal ★10 · 📚41 — Stop guessing thresholds. It fits per-question thresholds on your labels, verifies them on held-out data, and fails CI when a model update breaks them. Cited in more "best practice" sections than any other tool.
  • sutro-sh/jev-align ★300 · 📚42 — Builds calibrated "AI functions" from human feedback, using active labelling plus GEPA to optimize the questions.
  • AbdelStark/jev-benchmarks ★21 · 📚33 · jmanhype/jev-dspy-lab ★10 · 📚22 — Calibration, selective risk, and record/replay in DSPy.
  • openlayer-ai/jevals ★98 · 📚26 — Agent evals and guardrails as typed decisions (it also runs locally with Kev or Laya): eight judgments per trace for ~$0.00006.
  • smkrv/jev-calibrate ★31 · 📚17 · nikkoxgonzales/jev-certify ★1 · 📚11 · sathariels/jevcheck ★2 · 📚7 · vcjdeboer/jev-reliability 📚4 — Criteria tuning, conformal routing bounds, behavioural contract tests that refuse jev-latest, and repeatability/phrasing preflights.
  • suraj-phanindra/wellposed ★2 · 📚17 · ariel-frischer/jevkit ★3 · 📚22 · simota/tenbin ★4 · 📚18 — Lint your questions offline, before you pay for calls.
  • AntonioCoppe/jev-harness ★15 · 📚31 — Confidence gates, shadow mode, recipes and evals (on a row-filter job, Claude CLI took 48.9 s and the decision model 1.3 s).
  • mandu5/jevcompat ★0 · 📚6 — A conformance suite for /v1/systemone-compatible servers. Use it to check whether a Jev-like model really is a drop-in.
  • → catalog/bench.md

Games, robotics and simulation

The recipe for games and control: feed structured state (RAM → JSON, legal-move lists), never pixels. Let code own the physics, the route and the arithmetic, and let the decision model pick at branches. Keep a fast deterministic reflex layer that can veto it.

Apps and domain applications


Benchmarks and independent studies

Summarized with numbers in docs/evidence.md.

Research papers

~35 papers on Jev and other System One models appeared within two weeks, all unreplicated preprints. The full annotated list is in docs/papers.md.

Articles, analysis and critique

Analysis

News

Community discussion

Learning resources

Directories and other lists


About this meta-list

How it was built

  1. Pull. All 106 source repositories listed in sources.md were shallow-cloned on 2026-09-30. Their star counts and commits were recorded.
  2. Union. Every Markdown link in every file was extracted: 127k link occurrences, which de-duplicate to 18,330 unique entries after URL normalization. GitHub deep links collapse to owner/repo, and links to the source lists themselves are dropped.
  3. Consensus. For each entry, 📚 counts how many distinct lists cite it.
  4. Verification. The 2,533 GitHub entries cited by ≥5 lists were fetched live to confirm they exist, follow renames, and record stars and descriptions. The 21 that didn't resolve are listed in catalog/unresolved.md.
  5. Categorization. Entries were categorized automatically for the catalog. The sections above were hand-curated.
  6. Synthesis. Every source list was read in full. Per-source notes are in docs/source-notes/. The insights, the reconciled facts and the evidence digest in docs/ were synthesized from those notes.

Caveats

  • This is a snapshot of a two-week-old ecosystem. Specs, prices and access change quickly, so check the official models page.
  • Most numbers are author-reported and unreplicated. They are labelled as such where it matters.
  • Inclusion is not endorsement. Many launch-week repos are thin scaffolds, and many tools send your data to third parties.
  • This list is not affiliated with any model vendor.

Updating

scripts/refresh.sh      # re-pull all sources → stars/commits → catalog → sources.md → README numbers

To add a source, append its GitHub URL to sources.md and refresh. The hand-written synthesis in docs/ and templates/README.md.tmpl should then be reviewed against git diff catalog/.

Repository layout

README.md                  ← generated from templates/README.md.tmpl (live ★/📚 numbers)
sources.md                 ← the 106 sources: stars, pull date, commit
docs/                      ← synthesis: what-is-jev, patterns, evidence, open-models, papers, source-lists
docs/source-notes/         ← one review note per source list
catalog/                   ← the complete categorized union (18,330 entries)
data/                      ← catalog.json/csv, sources.json, verified.tsv, source_profiles.json
scripts/                   ← pull_sources, extract_links, build_catalog, verify_github, render_*, refresh.sh
awesome
awesome-list
awesome-lists
jev
jev-ai
jev-alternative
jev-api
jev-model
system-one
systemone

stefanwebb/meta-awesome-jev

The Awesome List of Awesome Jev Lists

Python

2

4 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

Meta Analysis of "Awesome Jev" GitHub Repos (r/LocalLLaMA)

I noticed that there are heaps of Awesome Jev resource list posts popping up on GitHub (e.g. https://github.com/yibie/awesome-jev) so I thought why not ask Claude to pull them all in and do a meta analysis of insights and applications. Sharing in case it's of interest:…

0

Oct 1, 2026

README

Meta Awesome Jev-like Models Awesome

A guide to the community of Jev-like models. These are "System One" decision models: you send them a state and some typed questions, and they return calibrated choices, scores and yes/no probabilities instead of generated text.

Jev introduced the /v1/systemone interface in September 2026. Within weeks it was served by a growing family of hosted and open models, including Mercury Decide, Liquid d1, Solar Decide, Laya, Kev, Decider and more.

This list merges, de-duplicates, cross-checks and summarizes 106 community "awesome-jev" repositories.

18,330 unique links from 106 lists · 5,410 cited by 3 or more lists · 2,533 GitHub repos verified live · sources last pulled 2026-09-30 (sources.md)

★ = GitHub stars when verified on the pull date · 📚N = the number of the 106 source lists that cite the entry. 📚 is the cross-list consensus signal: an entry that many independently curated lists include has been vetted many times.


Contents

Deep dives in this repo:

docs/open-models.mdThe Jev-like model family: open replications, local runtimes and hosted alternatives, with caveats
docs/patterns.mdQuestion design, composition patterns, thresholds, economics, security, a production checklist and anti-patterns
docs/evidence.mdEvery independent benchmark, calibration audit, robustness probe and negative result, with numbers
docs/what-is-jev.mdA reference for Jev, the original model: specs, API, timeline, and contradictions between sources, resolved
docs/papers.md~35 papers on Jev and Jev-like models, plus the research lineage
docs/source-lists.mdA review of all 106 source lists: which to read for what, and which to treat with caution
catalog/The complete union of every link from every list, in 21 categories, ranked by consensus

Jev-like models in 60 seconds

A Jev-like, or System One, model is not a chat model, and it never writes text. You send it two things in one request:

  • a state: the evidence (a string, a JSON object, or an array of text);
  • a map of typed questions about that state.

It answers every question in parallel, in a single call. Each answer is a probability distribution that your code can threshold.

Every System One model in the family speaks the same /v1/systemone wire format. Here is a support ticket with one question of each type:

POST /v1/systemone
{
  "model": "<model id>",
  "state": {
    "ticket": "I was charged twice for my order last week and nobody has answered my emails. I want my money back today."
  },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this ticket?",
      "criteria": {
        "billing":   "Charges, refunds, invoices",
        "technical": "Bugs, outages, errors",
        "other":     "Anything else"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Annoyed but civil", "Angry or threatening to leave"]
    },
    "wants_refund": {
      "type": "noul",
      "instructions": "Is the customer asking for a refund?"
    }
  }
}

The response contains one typed answer per question:

{
  "model": "<resolved model version>",
  "answers": {
    "team":         { "type": "choice", "choice": "billing",
                      "probabilities": { "billing": 0.95, "technical": 0.02, "other": 0.03 },
                      "confidence": 0.93 },
    "frustration":  { "type": "score", "score": 1.62,
                      "legend": { "0": "Calm", "1": "Annoyed but civil", "2": "Angry or threatening to leave" },
                      "probabilities": { "0": 0.03, "1": 0.32, "2": 0.65 }, "confidence": 0.71 },
    "wants_refund": { "type": "noul", "noul": 0.97 }
  },
  "usage": { "input_tokens": 214, "output_tokens": 5 }
}

Your code then decides what to do with the answers:

a = response["answers"]
if a["team"]["confidence"] >= 0.8:          # "the answer is what; confidence is whether to act"
    route_to(a["team"]["choice"])
else:
    send_to_human_triage()
if a["wants_refund"]["noul"] > 0.9 and a["frustration"]["score"] >= 1.5:
    escalate_priority()
Question typeAsksYou supplyYou get back
Choice"Which one?"an instructions string and criteria as a map of option → description (up to 255 options; always include an other)choice, probabilities per option, confidence
Score"Where on this ordered scale?"an instructions string and criteria as an ordered list of 2–10 levelsscore (probability-weighted and 0-indexed, so it can fall between levels), legend, probabilities, confidence
Noul"Is this true?"an instructions stringnoul: the probability, from 0 to 1, that the answer is yes. There is no confidence field, and 0.5 means "can't tell", not "medium".

→ The reference for the original model and the question-design guide go deeper.

Quick Start: Jev-like inference for free

This is how to use a Jev-like (System One) model for free. Mercury Decide is Inception's System One decision model, served on OpenRouter as inception/mercury-decide:free. Details from its OpenRouter page:

  • $0 input and $0 output. Free endpoints are rate-limited.
  • 32,768-token context.
  • About 0.42 s p50 latency and up to ~14 decisions per second.
  • The same /v1/systemone question schema shown above.

1. Get a free OpenRouter API key at openrouter.ai/settings/keys.

2. Call the Decisions API. Mercury Decide is a decisions model, so it uses OpenRouter's Decisions API, not /chat/completions. OpenAI-style chat SDKs won't work with it.

export OPENROUTER_API_KEY=sk-or-...

curl https://openrouter.ai/api/alpha/decisions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "inception/mercury-decide:free",
    "state": "I was charged twice for my order and want my money back.",
    "questions": {
      "team":   { "type": "choice", "instructions": "Which team should handle this ticket?",
                  "criteria": { "billing": "Charges and refunds", "technical": "Bugs and outages", "other": "Anything else" } },
      "refund": { "type": "noul",   "instructions": "Is the customer asking for a refund?" }
    }
  }'

3. Or call it from Python (pip install requests; no other SDK is needed):

import os, requests

resp = requests.post(
    "https://openrouter.ai/api/alpha/decisions",
    headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
    json={
        "model": "inception/mercury-decide:free",
        "state": {"ticket": "I was charged twice for my order and want my money back."},
        "questions": {
            "team": {"type": "choice", "instructions": "Which team should handle this ticket?",
                     "criteria": {"billing": "Charges and refunds", "technical": "Bugs and outages",
                                  "other": "Anything else"}},
            "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
            "urgency": {"type": "score", "instructions": "How urgent is this ticket?",
                        "criteria": ["Can wait a week", "Handle within a day", "Handle within the hour"]},
        },
    },
    timeout=30,
)
resp.raise_for_status()
answers = resp.json()["answers"]
print(answers["team"]["choice"], answers["team"]["confidence"], answers["refund"]["noul"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"])  # e.g. 1.05 {'0': 0.04, '1': 0.87, '2': 0.09}

Tips:

  • Already have code for another /v1/systemone client? OpenRouter also accepts the same request body at https://openrouter.ai/api/v1/systemone, so you can point an existing client's base URL at https://openrouter.ai/api.
  • Batch questions. Put every independent question about a state into one request. It costs the same as a single question, so use the batching to stay inside the free rate limit.
  • Graduate when you need more. Move to a paid decision model, or self-host an open one (Laya, Kev or Ollaya), when you need higher limits or local data. Most of these keep the same schema; see the model family.
  • Evaluate on your own labelled examples before trusting thresholds. Calibration doesn't transfer between models, so a 0.8 cut-off tuned on one model means nothing on another.

What 106 lists taught us: key insights

The same lessons show up again and again across the lists. Most measurements were taken on Jev, the first and most-tested model, and are labelled as such. The lessons about how to build apply to the whole family. Numbers link to the underlying studies in docs/evidence.md.

  1. Jev-like models are a System One decision layer, not a model swap. The winning architecture is always the same: code builds the candidates, the decision model picks or scores, code acts. An LLM is called only for text or open reasoning. Examples:
    • Browser agents let the model choose the element and use an LLM only to type.
    • Games give the model the legal moves.
    • Extraction lets regex find spans and the model select one.
  2. The 100× speedups are real only against chains of LLM calls. Jev's launch figures ("193.6× faster / 444.6× cheaper") are a best case against the slowest comparator; the average across eight setups is 97.8× / 149.2×. A single question is only ~1.5–3× faster than a small chat model. The order-of-magnitude wins come from batching many questions per call, where latency is flat up to ~25 questions and batching 13 questions was 12.2× cheaper, and from deleting LLM calls that never needed an LLM.
  3. Question design is the whole game.
    • Splitting one question into atomic Nouls took a task from 48% → 98%, and phishing detection from 62.6% → 95.0%.
    • Rewriting the criteria moved accuracy from 70% → 96%.
    • Renaming options from 0/1 to no/yes cut AUC from .81 to .58.
    • Always include an "other/unknown" option. Removing it took one benchmark from 0.95 to 0.00.
  4. Keep computation in code. Counting (33–65%), arithmetic, dates, multi-hop reasoning and open-ended generation are documented weak spots. Jev's own jaggedness page lists them, and independent tests on open replicas agree. Use one Noul per item and sum in code.
  5. Confidence is a routing signal, not a guarantee.
    • The top-confidence band is very reliable (48/48 correct at ≥ 0.9 in one audit).
    • "Confidently wrong" still happens, and Choice and Score skew overconfident.
    • Calibration does not transfer across datasets, languages, phrasings, model versions or models.
    • Fit per-question thresholds on your own labels, and pin a model version.
  6. Cascades are the most consistent cost/accuracy win, with one catch. Accepting confident decisions and escalating the rest to an LLM matched frontier accuracy at ~¼ the cost in several studies. But LLMs repeated ~96% of Jev's confident errors, so cascades mostly save money rather than add accuracy. Price the fallback: one production replay was 96% cheaper per call but 4% more expensive overall.
  7. It is not a security boundary. Typed output blocks format attacks, not decision manipulation. Blunt injected commands mostly fail. Evidence-shaped text flips decisions: fake approvals, editor's notes and fluent context flipped 61% of correct answers in JevOut, and 65–73% on open clones. Put deterministic rules first, keep humans on irreversible actions, and track taint.
  8. Audit what tools send. Hands-on audits found community tools leaking secrets, sending .pem files and screenshots, exposing keys, and failing open on API errors. Check data egress and fail-closed behaviour before installing.
  9. Coding agents are the dominant use case. The largest clusters across all lists are context compaction, model/effort routing, tool-call gates, "done" verification, and skill/rule pruning. Compaction is the most contested of these: several careful evaluations (Hermes Agent, jev-use) did not adopt it.
  10. Classical baselines are still strong. Trained small classifiers (bge-small + LR at 93% on Banking77, TF-IDF on spam) and regex often match or beat decision models. The durable edges of System One models are zero training, cost, latency, and robustness under distribution drift: on drifted spam, Jev held at 97.3% while TF-IDF fell to 72.5%.
  11. The family grew fast, but "compatible" ≠ "equivalent". Within days, dozens of open models and /v1/systemone servers appeared (Laya, Kev, SemIf, NanoJev, Decider, Ollaya), followed by hosted alternatives (Mercury Decide, Liquid d1, Solar Decide). The shared schema makes them drop-ins for your code, not for each other's accuracy or calibration. Most "beats Jev" claims are in-distribution, so measure on your own data.
  12. Most of the evidence is young. Nearly all numbers are author-reported, small-n, and from the family's first weeks, and the ~29 arXiv papers are unreplicated preprints. The best lists label every number (vendor / author-reported / independent), and this one does too.

Reference docs and SDKs

Jev's documentation is the most complete public description of the /v1/systemone interface: its question types, confidence semantics, composition patterns and cookbooks. Because Jev-like models share the schema, most of it applies to every System One model.

Interface docs

Reference SDKs and tools

  • Python: typesafe-sdk-python ★257 · 📚59. JS/TS: typesafe-sdk-js ★260 · 📚57. Both are /v1/systemone clients. Set the base-URL environment variable to point either one at any compatible server (local Kev, Decider or Ollaya; OpenRouter; and others).
  • system-one-adapter-python ★368 · 📚61 — The same client interface answered by an ordinary LLM. Use it for A/B baselines and fallback.
  • skills ★2,491 · 📚68 — An agent skill that teaches coding agents how to pick a question type and design questions.
  • For model-agnostic clients across many languages, see Community SDKs.

Patterns and cookbooks (the full annotated index covers all ~18)


Ecosystem

Hand-picked from the consensus of the source lists: mostly entries cited by many lists, plus a few high-signal ones that fewer lists noticed. Each section links to its full catalog page.

Jev-like models: hosted, open and local

The System One model family itself. Many members serve /v1/systemone, so the same client code works across them. docs/open-models.md maps the landscape and its caveats.

Hosted

  • Mercury Decide (Inception) — A structured decision model on OpenRouter, free, with a 32K context and the /v1/systemone schema. See the Quick Start.
  • Jev (TypeSafe) — The original and most-tested model, with a reference.
  • Liquid AI d1 — Serves /decisions/v1/systemone and has a free tier. Reportedly #1 on the Jev Decision Index.
  • Upstage Solar Decide — Same schema, 512K context.
  • Respan Span-01 — A System One model for behaviour monitoring.
  • The OpenAI Decisions API (Luna, preview).

Open and local

Framework and platform integrations

System One decision models landed inside mainstream frameworks within days of Jev's launch, usually as a new evaluate / decide / classify model type alongside generate. Most of these integrations target the shared /v1/systemone shape.

Community SDKs and clients

Every major language had a community SDK within a week. Check the last commit date before you depend on one.

MCP servers

Agent skills and plugins

Coding-agent tooling

The largest cluster in the ecosystem. It covers Claude Code, Codex, Pi, Hermes, OpenCode and Cursor.

Context compaction and pruning. This is contested, so read the evidence first.

Model and effort routing. Route at session boundaries and run in shadow mode first.

Guardrails and tool-call gates. Put deterministic rules first, and choose fail-open or fail-closed explicitly.

Done-verification, review and semantic lint

Browser, desktop and mobile

The shared recipe: build a numbered list of what's on the page or screen (accessibility tree, DOM or OCR), let the decision model pick the operation and target, let code execute, and call an LLM only when text must be typed. Most decision models are text-only and can't see pixels.

Search, RAG and data

Evaluation, calibration and QA tooling

  • abhixhek/jevcal ★10 · 📚41 — Stop guessing thresholds. It fits per-question thresholds on your labels, verifies them on held-out data, and fails CI when a model update breaks them. Cited in more "best practice" sections than any other tool.
  • sutro-sh/jev-align ★300 · 📚42 — Builds calibrated "AI functions" from human feedback, using active labelling plus GEPA to optimize the questions.
  • AbdelStark/jev-benchmarks ★21 · 📚33 · jmanhype/jev-dspy-lab ★10 · 📚22 — Calibration, selective risk, and record/replay in DSPy.
  • openlayer-ai/jevals ★98 · 📚26 — Agent evals and guardrails as typed decisions (it also runs locally with Kev or Laya): eight judgments per trace for ~$0.00006.
  • smkrv/jev-calibrate ★31 · 📚17 · nikkoxgonzales/jev-certify ★1 · 📚11 · sathariels/jevcheck ★2 · 📚7 · vcjdeboer/jev-reliability 📚4 — Criteria tuning, conformal routing bounds, behavioural contract tests that refuse jev-latest, and repeatability/phrasing preflights.
  • suraj-phanindra/wellposed ★2 · 📚17 · ariel-frischer/jevkit ★3 · 📚22 · simota/tenbin ★4 · 📚18 — Lint your questions offline, before you pay for calls.
  • AntonioCoppe/jev-harness ★15 · 📚31 — Confidence gates, shadow mode, recipes and evals (on a row-filter job, Claude CLI took 48.9 s and the decision model 1.3 s).
  • mandu5/jevcompat ★0 · 📚6 — A conformance suite for /v1/systemone-compatible servers. Use it to check whether a Jev-like model really is a drop-in.
  • → catalog/bench.md

Games, robotics and simulation

The recipe for games and control: feed structured state (RAM → JSON, legal-move lists), never pixels. Let code own the physics, the route and the arithmetic, and let the decision model pick at branches. Keep a fast deterministic reflex layer that can veto it.

Apps and domain applications


Benchmarks and independent studies

Summarized with numbers in docs/evidence.md.

Research papers

~35 papers on Jev and other System One models appeared within two weeks, all unreplicated preprints. The full annotated list is in docs/papers.md.

Articles, analysis and critique

Analysis

News

Community discussion

Learning resources

Directories and other lists


About this meta-list

How it was built

  1. Pull. All 106 source repositories listed in sources.md were shallow-cloned on 2026-09-30. Their star counts and commits were recorded.
  2. Union. Every Markdown link in every file was extracted: 127k link occurrences, which de-duplicate to 18,330 unique entries after URL normalization. GitHub deep links collapse to owner/repo, and links to the source lists themselves are dropped.
  3. Consensus. For each entry, 📚 counts how many distinct lists cite it.
  4. Verification. The 2,533 GitHub entries cited by ≥5 lists were fetched live to confirm they exist, follow renames, and record stars and descriptions. The 21 that didn't resolve are listed in catalog/unresolved.md.
  5. Categorization. Entries were categorized automatically for the catalog. The sections above were hand-curated.
  6. Synthesis. Every source list was read in full. Per-source notes are in docs/source-notes/. The insights, the reconciled facts and the evidence digest in docs/ were synthesized from those notes.

Caveats

  • This is a snapshot of a two-week-old ecosystem. Specs, prices and access change quickly, so check the official models page.
  • Most numbers are author-reported and unreplicated. They are labelled as such where it matters.
  • Inclusion is not endorsement. Many launch-week repos are thin scaffolds, and many tools send your data to third parties.
  • This list is not affiliated with any model vendor.

Updating

scripts/refresh.sh      # re-pull all sources → stars/commits → catalog → sources.md → README numbers

To add a source, append its GitHub URL to sources.md and refresh. The hand-written synthesis in docs/ and templates/README.md.tmpl should then be reviewed against git diff catalog/.

Repository layout

README.md                  ← generated from templates/README.md.tmpl (live ★/📚 numbers)
sources.md                 ← the 106 sources: stars, pull date, commit
docs/                      ← synthesis: what-is-jev, patterns, evidence, open-models, papers, source-lists
docs/source-notes/         ← one review note per source list
catalog/                   ← the complete categorized union (18,330 entries)
data/                      ← catalog.json/csv, sources.json, verified.tsv, source_profiles.json
scripts/                   ← pull_sources, extract_links, build_catalog, verify_github, render_*, refresh.sh
awesome
awesome-list
awesome-lists
jev
jev-ai
jev-alternative
jev-api
jev-model
system-one
systemone

Languages

Python

94.3%

Shell

5.7%