A guide to the community of Jev-like models. These are "System One" decision models: you send them a state and some typed questions, and they return calibrated choices, scores and yes/no probabilities instead of generated text.
Jev introduced the
/v1/systemoneinterface in September 2026. Within weeks it was served by a growing family of hosted and open models, including Mercury Decide, Liquid d1, Solar Decide, Laya, Kev, Decider and more.This list merges, de-duplicates, cross-checks and summarizes 106 community "awesome-jev" repositories.
18,330 unique links from 106 lists · 5,410 cited by 3 or more lists · 2,533 GitHub repos verified live · sources last pulled 2026-09-30 (sources.md)
★ = GitHub stars when verified on the pull date · 📚N = the number of the 106 source lists that cite the entry. 📚 is the cross-list consensus signal: an entry that many independently curated lists include has been vetted many times.
Deep dives in this repo:
| docs/open-models.md | The Jev-like model family: open replications, local runtimes and hosted alternatives, with caveats |
| docs/patterns.md | Question design, composition patterns, thresholds, economics, security, a production checklist and anti-patterns |
| docs/evidence.md | Every independent benchmark, calibration audit, robustness probe and negative result, with numbers |
| docs/what-is-jev.md | A reference for Jev, the original model: specs, API, timeline, and contradictions between sources, resolved |
| docs/papers.md | ~35 papers on Jev and Jev-like models, plus the research lineage |
| docs/source-lists.md | A review of all 106 source lists: which to read for what, and which to treat with caution |
| catalog/ | The complete union of every link from every list, in 21 categories, ranked by consensus |
A Jev-like, or System One, model is not a chat model, and it never writes text. You send it two things in one request:
It answers every question in parallel, in a single call. Each answer is a probability distribution that your code can threshold.
Every System One model in the family speaks the same /v1/systemone wire format. Here is a support ticket with one question of each type:
POST /v1/systemone
{
"model": "<model id>",
"state": {
"ticket": "I was charged twice for my order last week and nobody has answered my emails. I want my money back today."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "Bugs, outages, errors",
"other": "Anything else"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Annoyed but civil", "Angry or threatening to leave"]
},
"wants_refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}
The response contains one typed answer per question:
{
"model": "<resolved model version>",
"answers": {
"team": { "type": "choice", "choice": "billing",
"probabilities": { "billing": 0.95, "technical": 0.02, "other": 0.03 },
"confidence": 0.93 },
"frustration": { "type": "score", "score": 1.62,
"legend": { "0": "Calm", "1": "Annoyed but civil", "2": "Angry or threatening to leave" },
"probabilities": { "0": 0.03, "1": 0.32, "2": 0.65 }, "confidence": 0.71 },
"wants_refund": { "type": "noul", "noul": 0.97 }
},
"usage": { "input_tokens": 214, "output_tokens": 5 }
}
Your code then decides what to do with the answers:
a = response["answers"]
if a["team"]["confidence"] >= 0.8: # "the answer is what; confidence is whether to act"
route_to(a["team"]["choice"])
else:
send_to_human_triage()
if a["wants_refund"]["noul"] > 0.9 and a["frustration"]["score"] >= 1.5:
escalate_priority()
| Question type | Asks | You supply | You get back |
|---|---|---|---|
| Choice | "Which one?" | an instructions string and criteria as a map of option → description (up to 255 options; always include an other) | choice, probabilities per option, confidence |
| Score | "Where on this ordered scale?" | an instructions string and criteria as an ordered list of 2–10 levels | score (probability-weighted and 0-indexed, so it can fall between levels), legend, probabilities, confidence |
| Noul | "Is this true?" | an instructions string | noul: the probability, from 0 to 1, that the answer is yes. There is no confidence field, and 0.5 means "can't tell", not "medium". |
→ The reference for the original model and the question-design guide go deeper.
This is how to use a Jev-like (System One) model for free. Mercury Decide is Inception's System One decision model, served on OpenRouter as inception/mercury-decide:free. Details from its OpenRouter page:
/v1/systemone question schema shown above.1. Get a free OpenRouter API key at openrouter.ai/settings/keys.
2. Call the Decisions API. Mercury Decide is a decisions model, so it uses OpenRouter's Decisions API, not /chat/completions. OpenAI-style chat SDKs won't work with it.
export OPENROUTER_API_KEY=sk-or-...
curl https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inception/mercury-decide:free",
"state": "I was charged twice for my order and want my money back.",
"questions": {
"team": { "type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": { "billing": "Charges and refunds", "technical": "Bugs and outages", "other": "Anything else" } },
"refund": { "type": "noul", "instructions": "Is the customer asking for a refund?" }
}
}'
3. Or call it from Python (pip install requests; no other SDK is needed):
import os, requests
resp = requests.post(
"https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
json={
"model": "inception/mercury-decide:free",
"state": {"ticket": "I was charged twice for my order and want my money back."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Charges and refunds", "technical": "Bugs and outages",
"other": "Anything else"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["Can wait a week", "Handle within a day", "Handle within the hour"]},
},
},
timeout=30,
)
resp.raise_for_status()
answers = resp.json()["answers"]
print(answers["team"]["choice"], answers["team"]["confidence"], answers["refund"]["noul"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"]) # e.g. 1.05 {'0': 0.04, '1': 0.87, '2': 0.09}
Tips:
/v1/systemone client? OpenRouter also accepts the same request body at https://openrouter.ai/api/v1/systemone, so you can point an existing client's base URL at https://openrouter.ai/api.The same lessons show up again and again across the lists. Most measurements were taken on Jev, the first and most-tested model, and are labelled as such. The lessons about how to build apply to the whole family. Numbers link to the underlying studies in docs/evidence.md.
0/1 to no/yes cut AUC from .81 to .58..pem files and screenshots, exposing keys, and failing open on API errors. Check data egress and fail-closed behaviour before installing./v1/systemone servers appeared (Laya, Kev, SemIf, NanoJev, Decider, Ollaya), followed by hosted alternatives (Mercury Decide, Liquid d1, Solar Decide). The shared schema makes them drop-ins for your code, not for each other's accuracy or calibration. Most "beats Jev" claims are in-distribution, so measure on your own data.Jev's documentation is the most complete public description of the /v1/systemone interface: its question types, confidence semantics, composition patterns and cookbooks. Because Jev-like models share the schema, most of it applies to every System One model.
Interface docs
📚31 — POST /v1/systemone, with request and answer shapes for all three primitives. The full docs index (llms.txt) 📚22 is also available.📚27📚30📚10📚13📚17📚17📚35 — Read this before building. Nine documented failure modes (literal reading, math, dates, indirection, distracting state, adversarial content and more). They are written for Jev but are a good test plan for any Jev-like model.📚57 — The launch post that defined the category: RLCD (calibrated-decision training), the parallel sampler, and the Doom and Wikiracing demos.📚32 — Vendor evaluations on four workflows. The reference labels are averaged from two frontier models, not ground truth. The code is in WorkflowEvals 📚3.Reference SDKs and tools
★257 · 📚59. JS/TS: typesafe-sdk-js ★260 · 📚57. Both are /v1/systemone clients. Set the base-URL environment variable to point either one at any compatible server (local Kev, Decider or Ollaya; OpenRouter; and others).★368 · 📚61 — The same client interface answered by an ordinary LLM. Use it for A/B baselines and fallback.★2,491 · 📚68 — An agent skill that teaches coding agents how to pick a question type and design questions.Patterns and cookbooks (the full annotated index covers all ~18)
📚19:
📚16📚15📚14📚14📚23 — 13 questions in one call: 12.2× cheaper and 10× faster.📚15 — Beam search past the 255-option cap.📚15 — Legal search: top-1 5% → 18%.📚16📚17📚18📚14📚16📚13📚17📚17Hand-picked from the consensus of the source lists: mostly entries cited by many lists, plus a few high-signal ones that fewer lists noticed. Each section links to its full catalog page.
The System One model family itself. Many members serve /v1/systemone, so the same client code works across them. docs/open-models.md maps the landscape and its caveats.
Hosted
/v1/systemone schema. See the Quick Start./decisions/v1/systemone and has a free tier. Reportedly #1 on the Jev Decision Index.Open and local
★29,237 · 📚38 — Convai's Apache-2.0 non-autoregressive encoder (421M/322M, 100+ languages, ~33 ms). Runtimes: laya-mlx ★6,654 · 📚23 (7–14 ms on M3 Max) and receptron/laya ★661 · 📚20 (Node/ONNX).★8,057 · 📚57 — A trainable family of System One models on Qwen3.5/3.8 with a pointer head and a /v1/systemone server.★4,615 · 📚63 — "Semantic ifs" from open models on a 3090, using shared-prefix logit readout.★2,453 · 📚60 · vinnylarouge/jevlike ★1,335 · 📚57 · Mapika/decider ★993 · 📚45 · bespokelabsai/nimble ★1,981 · 📚21 · wfzyx/von ★794 · 📚48 — Trained replicas and recipes (0.4B–35B), several with published limits.★986 · 📚32 · featherless-ai/simple-jev ★575 · 📚40 · razorback16/openjev ★548 · 📚49 · ekzhang/openjev-sglang ★335 · 📚48 · githubnext/localjev ★803 · 📚16 — Turn any open LLM into a decision endpoint, with no training. AnyJev adds option-order correction.★1,032 · 📚24 — "Ollama for decision models": it runs open System One models locally. It serves Laya, Decider, NLI and GLiClass behind a /v1/systemone-compatible API.System One decision models landed inside mainstream frameworks within days of Jev's launch, usually as a new evaluate / decide / classify model type alongside generate. Most of these integrations target the shared /v1/systemone shape.
★27,056 · 📚7 — The AI SDK's decision-model provider maps Choice, Score and Boolean onto experimental_evaluate.★5,426 · 📚28 — Vercel's agent framework. A decision model (Jev by default) powers its evaluate path and its tool-approval policies.★3,234 · 📚9 · vercel-labs/ai-cli ★817 · 📚23 — fx has a decision-model permission reviewer, reported at p95 18× faster than GPT Luna. ai-cli adds an evaluate command.★20,296 · 📚11 — A decision-model backend: output_type fields become typed questions (bool → Noul, Literal → Choice, IntEnum → Score).★59,945 · 📚9 — A decision-model complexity router, a compaction guardrail, and a pass-through proxy.★30,374 · 📚14 — A decision-model provider for tool and argument selection.★9,363 · 📚5 · BoundaryML/feelings ★22 · 📚10 — Maps the type system onto the primitives. feelings is a typed .feels() "AI if-statement".★1,240 · 📚7 — Records decision-model calls (Choice/Score/Noul) as OpenTelemetry spans. See also the Langfuse integration.★27,561 · 📚17 — A computer-use platform with a bounded-action recipe for decision models.★5,109 · 📚12 — A decision-model guardrail for jailbreaks, harmful content and secret leakage at the gateway.★2,445 · 📚16 — Nine search notebooks that use a decision model for ranking, filtering, routing and stopping.★8,206 · 📚7 — Uses a Noul to check whether a cached answer serves a new request.★41 · 📚20 · crmne/ruby_llm ★4,423 · 📚5 · laravel/ai ★1,203 · 📚6 — Java/Spring, Ruby and Laravel support.★143 · 📚23 — Makes decide a first-class inference type next to generate and stream, across 40 providers.★250,326 · 📚5 — A decision-model provider, skill routing, and a published (negative) compaction scorecard.📚3 — A production link-safety gate on typesafe-ai/jev.decide()), DeepEval, MotherDuck prompt_jev(), OpenRouter cookbook: gate tool calls, Ollama 0.35 decision models · → catalog/sdk.mdEvery major language had a community SDK within a week. Check the last commit date before you depend on one.
★6 · 📚30 — supports several gateways.★5 · 📚29★39 · 📚18 — SDK and CLI.★9 · 📚24★13 · 📚33 — async and blocking clients, observable retries.★6 · 📚31★1 · 📚23★0 · 📚6★51 · 📚32★7 · 📚35★18 · 📚30★16 · 📚24 — probabilistic control flow for Rails.★34 · 📚43 — pattern-match decisions inside OTP.★5 · 📚36★12 · 📚34 · Hawxy/TypeSafeAI.Net ★4 · 📚24★10 · 📚23★5 · 📚29★4 · 📚25★16 · 📚22★54 · 📚20 — maps Apple Foundation Models @Generable types onto the primitives.★7 · 📚20★96 · 📚48★2 · 📚30★39 · 📚8★20 · 📚18 — pandas and Polars.★6 · 📚27★8 · 📚23 — semantic rules inside Zod schemas.★3 · 📚22★25 · 📚39★13 · 📚32 — a CLI and stdio MCP server.★75 · 📚37 — Unix-pipeline and CI decisions with exit codes.★10 · 📚11 — jq meets decision models.★469 · 📚66 — The most-cited MCP server. Its typed tools (verify, screen for injection, find, rerank, classify, decide) follow a fail-closed contract. Install with claude mcp add jev -- npx -y @jkudish/jev-mcp.★337 · 📚61 — Formerly typesafe-mcp. A single connector for Jev, d1, CLM and Laya.★25 · 📚42 · burnigtm/jev-mcp ★61 · 📚29 · BYK/jev-mcp ★3 · 📚20 · PyModel/jev-judge-mcp ★69 · 📚17 · Brainwires/jevwire ★21 · 📚28 — Alternatives. BYK's server is eval-first, and jevwire adds an escalate-only Claude Code plugin.★47 · 📚18 — Classify first, read selectively. Batch text classification for agents.★18 · 📚17 — Hands off a whole browser task in one MCP call.★21 · 📚21 — An MCP proxy that screens both tool calls and tool results.jev-mcp repos share a name, so check the owner. → catalog/mcp.md★145 · 📚39 — The best community skill for writing decision-model programs: pick the primitive your code branches on, split two-property questions, and batch everything that shares a state.★932 · 📚42 — The most complete decision-model-in-the-agent-loop pack: routing, memory, compaction, skill selection, and computer and browser use (Hermes, Claude Code, Codex).★554 · source list — Five skills, a jev-decide CLI, 108 agent-supervision scenarios, and honest evals, including a negative agent result.★125 · 📚49 · ShivamPansuriya/jev-skill-gate ★7 · 📚24 · GodsBoy/jev-agent-skill-router ★23 · 📚33 · EliaAlberti/jev-rules ★62 · 📚27 — Show the agent only the skills and rules the current turn needs, and allow "none". jev-skill-gate cuts the skill manifest by ~75%.★33 · 📚34 — Hands agent steps that need no text output to a System One model (p50 ~230 ms). Its compaction and agreement evals are published.★36 · 📚30 · samtay32/jev-system-architect ★2 · 📚23 — Audit a codebase for fuzzy judgments that could become typed decision points.★2 · source list · Pleo2/awesome-jev-agent-skills ★0 · source list — Ready-made proceed/ask/stop gates, triage, diff review and QA-evidence skills.The largest cluster in the ecosystem. It covers Claude Code, Codex, Pi, Hermes, OpenCode and Cursor.
Context compaction and pruning. This is contested, so read the evidence first.
★7,250 · 📚71 — The most-starred project in the category. It replaces Claude Code's compaction summary with keep/drop decisions per tool call, and kept content stays verbatim. It began as a viral demo.★153 · 📚35 — Trims long Bash output before the model sees it.★99 · 📚39 — A calibrated context sieve with stub/recall pointers. It falls back to an LLM-backed adapter when the decision model is unreachable.★189 · 📚15 — Asks "is this a safe moment to compact?"★13 · 📚27 · leonaaardob/fast-dev-compaction ★9 · 📚26 · compozy/yoshi ★27 · 📚30 — Ports to Pi and Codex, and a pruning proxy.Model and effort routing. Route at session boundaries and run in shadow mode first.
★505 · 📚60 — Routes Claude Code and Codex to the cheapest capable model.★277 · 📚54 — Per-turn model, reasoning-effort and speed routing for Codex (~60% savings on a 237-turn backtest).★309 · 📚42 — One candidate set covering models, tools, subagents and skills.★16 · 📚21 · adarshmishra07/jcm-router ★8 · 📚27 · xinyao27/jevonian ★15 · 📚26 — Cache-aware routing that leaves the cached main chat alone.★293 · 📚15 · nidhi-singh02/agent-router ★99 · 📚25 · prismhq/jev-router ★14 · 📚29 — Mid-run effort selection, agent/CLI selection, and a router on LiteLLM.Guardrails and tool-call gates. Put deterministic rules first, and choose fail-open or fail-closed explicitly.
★48 · 📚38 — Deny/ask/allow auto mode for any coding agent. It uses session context, flags injection in tool results, and tracks taint. Smoke test.★153 · 📚50 — Steers instead of interrupting. Rule breaks fell from 6 to 0 over 150 paired runs, and a replay on 17k calls held 42.★618 · 📚61 — A supervisor for Codex and OpenCode workers, with risk-tiered thresholds.★30 · 📚38 — Auto-approves Bash, write and edit calls, and fails closed.★150 · 📚54 · Nyarlathoteppppp/pi-heed ★10 · 📚27 · harshwasan/jev-sentinel ★12 · 📚19 — Pi extensions that judge whether an action is destructive, exfiltrates data, stays in scope, or matches what you asked for.★5 · 📚19 · jesset/pi-verdict ★10 · 📚14 — Rules first, then a decision model for the grey zone. jev-engineering has a rerunnable 300-call injection test.★19 · 📚27 — Hermes command approvals: 8.7× faster, with 4.4× fewer prompts across 153 real commands.★32 · 📚27 · hemanth/pkg-gate ★0 · 📚16 — Screen codebases and npm lifecycle scripts before you run them.Done-verification, review and semantic lint
★646 · 📚74 — Staged code review (risk → file profile → evidence → severity → reviewer routing) with a local dashboard.★19 · 📚41 · qkal/Canny ★106 · 📚32 · noplan-inc/limpet ★5 · 📚24 — Stop-hook gates that block an unverified "done". Canny's rule: "deterministic hooks decide, the model advises".★469 · 📚19 — Enforces AGENTS.md rules with one Noul per rule.★12 · 📚39 — A pre-commit check that the message matches the diff; it blocks only on leaked credentials.★230 · 📚48 · supercorp-ai/supercov ★141 · 📚39 · lakeday-org/perch ★316 · 📚29 · mizchi/jev-lint ★113 · 📚23 · doeixd/jev-pref ★10 · 📚32 — Continuous quality review, coverage, and linting for semantic rules that a parser can't check.★2 · 📚19 · yamadashy/jev-labeler-action ★0 · 📚5 — GitHub Actions for submission review and issue/PR labels.The shared recipe: build a numbered list of what's on the page or screen (accessibility tree, DOM or OCR), let the decision model pick the operation and target, let code execute, and call an LLM only when text must be typed. Most decision models are text-only and can't see pixels.
★21,563 · 📚76 — The most-cited project in the entire ecosystem. It picks the operation and element in one request and booked a Zürich→London flight search in 7.1 s for $0.0039. The launch demo has been retested: 1/20 on complex tasks.★1,098 · 📚46 — macOS computer use: OCR the screen, then a decision model picks the action, for ~$0.0002 per step.★426 · 📚59 — Android automation; an Uber booking took 9 actions and ~21 s.★374 · 📚49 — Voice control at ~300 ms per spoken word. URLs are copied from your words, never generated.★297 · 📚49 · Ying-Kai-Liao/jev-browser ★93 · 📚43 · wy-coliney/jev-browser-use ★721 · 📚45 — In these, the LLM plans and the decision model decides. Ying-Kai-Liao's version solved 40/42 live tasks using ~70× fewer tokens.★613 · 📚28 · lahfir/agent-desktop ★1,733 · 📚16 · savka777/jev-use ★111 · 📚22 — Desktop control through accessibility trees, with local policy gates on sensitive clicks.★343 · 📚49 · realZachi/typesafe-adblock ★87 · 📚46 · anishfn/shapeshift ★755 · 📚23 — Browser extensions and UI: clutter removal, "is this DOM element an ad?", and a text box that morphs into the UI you mean.★495 · 📚64 — Web search where a decision model does source selection, query understanding and relevance ranking. It returns links, not generated answers.★386 · 📚59 — Ask Postgres tables questions in plain language. Its batching study found 20 rows per request gave 100%, and 80 rows gave 77–94%.★14 · 📚43 · colliber/duckdb-jev ★27 · 📚19 · mattn/sqlite3-jev ★3 · 📚15 · giuliosmall/pg_typesafe ★87 · 📚34 — Semantic SQL predicates. Every row leaves your machine, so pre-filter with cheap predicates first.★153 · 📚52 — Graph navigation: a Choice picks the next relationship and a Noul asks "goal reached?", with beam search in the app.★1,902 · 📚23 · keltokhy/jgrep ★131 · 📚32 · uehaj/sys1grep ★144 · 📚36 · sufianetaouil/every ★7 · 📚32 · ellipsis-dev/blink ★92 · 📚45 — "grep by meaning". These send code to the API, so check what leaves your machine.★487 · 📚31 — Fast document classification and packet splitting, from LlamaIndex's founder.★92 · 📚45 · RenaGao/jev-dataops ★60 · 📚18 — Dataset sifting (Rust, Parquet/JSONL) and streaming data selection with LoRA training.★8 · 📚33 · hev/reranker ★14 · 📚25 · hotchpotch/jev-reranker ★36 · 📚21 — Rerankers. Fuse with BM25 or embeddings rather than reranking with the decision model alone.★24 · 📚24 · keltokhy/jlink ★6 · 📚21 · keltokhy/jselect ★3 · 📚16 — Sort by meaning (pairwise comparisons plus Bradley-Terry), record linkage, and budgeted evidence selection.★16 · 📚33 · hyperspaceai/jevcache ★76 · 📚17 — OpenTelemetry log triage, and a decision cache for deterministic CI replay.★10 · 📚41 — Stop guessing thresholds. It fits per-question thresholds on your labels, verifies them on held-out data, and fails CI when a model update breaks them. Cited in more "best practice" sections than any other tool.★300 · 📚42 — Builds calibrated "AI functions" from human feedback, using active labelling plus GEPA to optimize the questions.★21 · 📚33 · jmanhype/jev-dspy-lab ★10 · 📚22 — Calibration, selective risk, and record/replay in DSPy.★98 · 📚26 — Agent evals and guardrails as typed decisions (it also runs locally with Kev or Laya): eight judgments per trace for ~$0.00006.★31 · 📚17 · nikkoxgonzales/jev-certify ★1 · 📚11 · sathariels/jevcheck ★2 · 📚7 · vcjdeboer/jev-reliability 📚4 — Criteria tuning, conformal routing bounds, behavioural contract tests that refuse jev-latest, and repeatability/phrasing preflights.★2 · 📚17 · ariel-frischer/jevkit ★3 · 📚22 · simota/tenbin ★4 · 📚18 — Lint your questions offline, before you pay for calls.★15 · 📚31 — Confidence gates, shadow mode, recipes and evals (on a row-filter job, Claude CLI took 48.9 s and the decision model 1.3 s).★0 · 📚6 — A conformance suite for /v1/systemone-compatible servers. Use it to check whether a Jev-like model really is a drop-in.The recipe for games and control: feed structured state (RAM → JSON, legal-move lists), never pixels. Let code own the physics, the route and the arithmetic, and let the decision model pick at branches. Keep a fast deterministic reflex layer that can veto it.
★422 · 📚64 — Super Mario Bros. from structured emulator RAM. It is the canonical "state as data" example.★228 · 📚66 — A MuJoCo drone. The decision model advises at 2.5 Hz while a 50 Hz reflex layer holds a veto.★9 · 📚38 — Code owns the route and the arithmetic, the model picks at branches in ~100 ms, and calibration is measured with Brier scores.★397 · 📚38 — 22 latency-focused demo apps by Nader Dabit.★27 · 📚45 · sorrycc/typesafe-snake ★23 · 📚32 · standardagents/jevpilot ★204 · 📚32 · rmalde/minecraft-agent ★568 · 📚18 — StarCraft, Snake (legal moves generated in code), a driving autopilot, and Minecraft (an LLM plans, the decision model acts).★249 · 📚30 · Dimweaker/jev-libero ★78 · 📚23 · TarunTomar122/jev-askable-arm ★11 · 📚27 · openroboto-ai/jev-robot-control ★53 · 📚13 — Robot arms that pick primitives from menus, never torques. OpenRoboto compares Jev with GPT-6 Astra on cost and time.★68 · 📚56 — Home Assistant: ask your house a question and get a number back, as sensors and automation actions.★2,708 · 📚71 — One trade decision per Monad block (~81 ms). It's educational: no list reports sustained trading profit, and bots should be dry-run by default.★483 · 📚41 — IRS form pages: 100% strict on 261 forms at ~$0.001 per page, replacing production Sonnet.★92 · 📚27 · parth-kp/jev-mail-classifier ★19 · 📚19 — Gmail triage: 1,000 emails in ~1 minute for ~3 cents.★7,181 · 📚37 — A read-only chat copilot for phones. An LLM drafts replies, a decision model ranks them, and you choose whether to send.★130 · 📚42 · AkashPriyadarshii/jev-seo ★88 · 📚32 · usenotra/notra ★225 · 📚19 — Social research, SEO/GEO audits, and a production GEO platform.★47 · 📚36 · backmeupplz/jev_antispam_bot ★13 · 📚20 — Discord and Telegram moderation.★12,344 · 📚29 — An AI trading OS that uses a decision model as an evidence and risk gate. It blocks entries but lets exits and protective orders through.★37 · 📚22 · youkiti/tiab-review-plugin — Systematic-review screening (95% recall on 16,645 records) and extraction.★244 · 📚33 · ChetasLua/jevmeter ★102 · 📚45 — Fun ones: kill, fix or ship your startup idea, and a live decision meter on any video.Summarized with numbers in docs/evidence.md.
★189 · 📚40 — JevBench: 534 frozen, half-sealed decisions and a four-axis composite across 40+ systems. See also the Benchmark Heaven leaderboard.📚12 — Jev vs six LLMs on human-labelled sets, with the data published.★0 · 📚11 — 123,805 pre-registered requests. ECE was 0.075 in-domain, but calibration fails OOD and authority injection works.★7 · 📚32 — The canonical question-atomization lesson: 62.6% with one question vs 95.0% atomized. Haiku 4.5 wins the direct question.★5 · 📚12 · bitnovus/jev-spam-eval ★2 · 📚34 — Classical baselines win in-domain, but Jev wins under drift.★9 · 📚39 · zhuyansen/jev-search-rerank-eval ★9 · 📚31 — Reranking: Jev ties Cohere, but doesn't beat embeddings alone, so fuse them.★0 · 📚15 · KantaHayashiAI/jev-does-not-play-dice ★3 · 📚17 · scienthoon/jev-ood-calibration ★6 · 📚25 · jourdanlabs/assay-001 ★0 · 📚18 — Calibration audits covering abstention, known probabilities, OOD, and a pre-registered test.★3 · 📚31 · zkousama/jagged 📚4 — Prompt injection and vulnerable-code detection.★86 · 📚18 — WebMCP: Jev + Mercury solved 49/49 tasks, bare Jev 25/49.★26 · 📚27 · RINNECODER/jev-behavior-study ★3 · 📚23 · FirasSX914/Janus ★2 · 📚25 — Where Jev holds up vs breaks, order and lost-in-the-middle effects, and when routing is worth it.📚21: one judge call vs dimension scores.📚19📚14~35 papers on Jev and other System One models appeared within two weeks, all unreplicated preprints. The full annotated list is in docs/papers.md.
📚18 — An ecosystem survey of 2,170 GitHub projects.📚7 · JEV vs. LLMs as Rubric Judges 📚5 — Jev comes within ~3 pts of GPT-6 at ~0.36% of the cost, but LLMs repeat its confident errors, which caps cascade gains.📚5 · JevOut 📚4 · JevAdvBench 📚5 · Decision Hijacking 📚6 — Robustness: option names, natural context and injection.📚4 · Jev-Mobile 📚4 · JEV-Star · Jev-Mem 📚6 — Planner/executor agents that make 66–73% fewer strong-model calls.📚4 · Just Ask Jev 📚7 — Annotation, large-scale coding and alignment-failure detection.Analysis
📚11 and 6 clones of Jev in 2 days.📚17 — Often named as the best long-form introduction.📚8 · TrueFoundry: What actually shipped · Silverthread Labs: What the 193× benchmark doesn't measure · ContextOS: production engineering review — Critiques.News
Community discussion
★124 · 📚10 — From one smart if-statement to a coding agent that reaches for a decision model on its own.★14 · 📚15 — A free hands-on course: 13 agent use cases, a fast brain and a slow brain, one OpenRouter key.★34 · 📚27 · datawhalechina/jev-cookbook ★46 · 📚14 · paramjeetn/jev-cookbook ★7 · 📚11 — Tested recipes on OpenRouter, 18 Chinese notebook recipes with evals and local fine-tuning, and 120+ use cases.★33 · 📚14 · Foadsf/jev-for-engineers ★5 · 📚23 · PromptEngineer48/langchain-jev-tutorial — An illustrated primitives guide, engineering examples, and a LangChain support-ops agent.📚9📚14: builds with reported cost and speed.📚4How it was built
owner/repo, and links to the source lists themselves are dropped.📚 counts how many distinct lists cite it.Caveats
Updating
scripts/refresh.sh # re-pull all sources → stars/commits → catalog → sources.md → README numbers
To add a source, append its GitHub URL to sources.md and refresh. The hand-written synthesis in docs/ and templates/README.md.tmpl should then be reviewed against git diff catalog/.
Repository layout
README.md ← generated from templates/README.md.tmpl (live ★/📚 numbers)
sources.md ← the 106 sources: stars, pull date, commit
docs/ ← synthesis: what-is-jev, patterns, evidence, open-models, papers, source-lists
docs/source-notes/ ← one review note per source list
catalog/ ← the complete categorized union (18,330 entries)
data/ ← catalog.json/csv, sources.json, verified.tsv, source_profiles.json
scripts/ ← pull_sources, extract_links, build_catalog, verify_github, render_*, refresh.sh
Python
94.3%
Shell
5.7%
A guide to the community of Jev-like models. These are "System One" decision models: you send them a state and some typed questions, and they return calibrated choices, scores and yes/no probabilities instead of generated text.
Jev introduced the
/v1/systemoneinterface in September 2026. Within weeks it was served by a growing family of hosted and open models, including Mercury Decide, Liquid d1, Solar Decide, Laya, Kev, Decider and more.This list merges, de-duplicates, cross-checks and summarizes 106 community "awesome-jev" repositories.
18,330 unique links from 106 lists · 5,410 cited by 3 or more lists · 2,533 GitHub repos verified live · sources last pulled 2026-09-30 (sources.md)
★ = GitHub stars when verified on the pull date · 📚N = the number of the 106 source lists that cite the entry. 📚 is the cross-list consensus signal: an entry that many independently curated lists include has been vetted many times.
Deep dives in this repo:
| docs/open-models.md | The Jev-like model family: open replications, local runtimes and hosted alternatives, with caveats |
| docs/patterns.md | Question design, composition patterns, thresholds, economics, security, a production checklist and anti-patterns |
| docs/evidence.md | Every independent benchmark, calibration audit, robustness probe and negative result, with numbers |
| docs/what-is-jev.md | A reference for Jev, the original model: specs, API, timeline, and contradictions between sources, resolved |
| docs/papers.md | ~35 papers on Jev and Jev-like models, plus the research lineage |
| docs/source-lists.md | A review of all 106 source lists: which to read for what, and which to treat with caution |
| catalog/ | The complete union of every link from every list, in 21 categories, ranked by consensus |
A Jev-like, or System One, model is not a chat model, and it never writes text. You send it two things in one request:
It answers every question in parallel, in a single call. Each answer is a probability distribution that your code can threshold.
Every System One model in the family speaks the same /v1/systemone wire format. Here is a support ticket with one question of each type:
POST /v1/systemone
{
"model": "<model id>",
"state": {
"ticket": "I was charged twice for my order last week and nobody has answered my emails. I want my money back today."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "Bugs, outages, errors",
"other": "Anything else"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Annoyed but civil", "Angry or threatening to leave"]
},
"wants_refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}
The response contains one typed answer per question:
{
"model": "<resolved model version>",
"answers": {
"team": { "type": "choice", "choice": "billing",
"probabilities": { "billing": 0.95, "technical": 0.02, "other": 0.03 },
"confidence": 0.93 },
"frustration": { "type": "score", "score": 1.62,
"legend": { "0": "Calm", "1": "Annoyed but civil", "2": "Angry or threatening to leave" },
"probabilities": { "0": 0.03, "1": 0.32, "2": 0.65 }, "confidence": 0.71 },
"wants_refund": { "type": "noul", "noul": 0.97 }
},
"usage": { "input_tokens": 214, "output_tokens": 5 }
}
Your code then decides what to do with the answers:
a = response["answers"]
if a["team"]["confidence"] >= 0.8: # "the answer is what; confidence is whether to act"
route_to(a["team"]["choice"])
else:
send_to_human_triage()
if a["wants_refund"]["noul"] > 0.9 and a["frustration"]["score"] >= 1.5:
escalate_priority()
| Question type | Asks | You supply | You get back |
|---|---|---|---|
| Choice | "Which one?" | an instructions string and criteria as a map of option → description (up to 255 options; always include an other) | choice, probabilities per option, confidence |
| Score | "Where on this ordered scale?" | an instructions string and criteria as an ordered list of 2–10 levels | score (probability-weighted and 0-indexed, so it can fall between levels), legend, probabilities, confidence |
| Noul | "Is this true?" | an instructions string | noul: the probability, from 0 to 1, that the answer is yes. There is no confidence field, and 0.5 means "can't tell", not "medium". |
→ The reference for the original model and the question-design guide go deeper.
This is how to use a Jev-like (System One) model for free. Mercury Decide is Inception's System One decision model, served on OpenRouter as inception/mercury-decide:free. Details from its OpenRouter page:
/v1/systemone question schema shown above.1. Get a free OpenRouter API key at openrouter.ai/settings/keys.
2. Call the Decisions API. Mercury Decide is a decisions model, so it uses OpenRouter's Decisions API, not /chat/completions. OpenAI-style chat SDKs won't work with it.
export OPENROUTER_API_KEY=sk-or-...
curl https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inception/mercury-decide:free",
"state": "I was charged twice for my order and want my money back.",
"questions": {
"team": { "type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": { "billing": "Charges and refunds", "technical": "Bugs and outages", "other": "Anything else" } },
"refund": { "type": "noul", "instructions": "Is the customer asking for a refund?" }
}
}'
3. Or call it from Python (pip install requests; no other SDK is needed):
import os, requests
resp = requests.post(
"https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": f"Bearer {os.environ['OPENROUTER_API_KEY']}"},
json={
"model": "inception/mercury-decide:free",
"state": {"ticket": "I was charged twice for my order and want my money back."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Charges and refunds", "technical": "Bugs and outages",
"other": "Anything else"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["Can wait a week", "Handle within a day", "Handle within the hour"]},
},
},
timeout=30,
)
resp.raise_for_status()
answers = resp.json()["answers"]
print(answers["team"]["choice"], answers["team"]["confidence"], answers["refund"]["noul"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"]) # e.g. 1.05 {'0': 0.04, '1': 0.87, '2': 0.09}
Tips:
/v1/systemone client? OpenRouter also accepts the same request body at https://openrouter.ai/api/v1/systemone, so you can point an existing client's base URL at https://openrouter.ai/api.The same lessons show up again and again across the lists. Most measurements were taken on Jev, the first and most-tested model, and are labelled as such. The lessons about how to build apply to the whole family. Numbers link to the underlying studies in docs/evidence.md.
0/1 to no/yes cut AUC from .81 to .58..pem files and screenshots, exposing keys, and failing open on API errors. Check data egress and fail-closed behaviour before installing./v1/systemone servers appeared (Laya, Kev, SemIf, NanoJev, Decider, Ollaya), followed by hosted alternatives (Mercury Decide, Liquid d1, Solar Decide). The shared schema makes them drop-ins for your code, not for each other's accuracy or calibration. Most "beats Jev" claims are in-distribution, so measure on your own data.Jev's documentation is the most complete public description of the /v1/systemone interface: its question types, confidence semantics, composition patterns and cookbooks. Because Jev-like models share the schema, most of it applies to every System One model.
Interface docs
📚31 — POST /v1/systemone, with request and answer shapes for all three primitives. The full docs index (llms.txt) 📚22 is also available.📚27📚30📚10📚13📚17📚17📚35 — Read this before building. Nine documented failure modes (literal reading, math, dates, indirection, distracting state, adversarial content and more). They are written for Jev but are a good test plan for any Jev-like model.📚57 — The launch post that defined the category: RLCD (calibrated-decision training), the parallel sampler, and the Doom and Wikiracing demos.📚32 — Vendor evaluations on four workflows. The reference labels are averaged from two frontier models, not ground truth. The code is in WorkflowEvals 📚3.Reference SDKs and tools
★257 · 📚59. JS/TS: typesafe-sdk-js ★260 · 📚57. Both are /v1/systemone clients. Set the base-URL environment variable to point either one at any compatible server (local Kev, Decider or Ollaya; OpenRouter; and others).★368 · 📚61 — The same client interface answered by an ordinary LLM. Use it for A/B baselines and fallback.★2,491 · 📚68 — An agent skill that teaches coding agents how to pick a question type and design questions.Patterns and cookbooks (the full annotated index covers all ~18)
📚19:
📚16📚15📚14📚14📚23 — 13 questions in one call: 12.2× cheaper and 10× faster.📚15 — Beam search past the 255-option cap.📚15 — Legal search: top-1 5% → 18%.📚16📚17📚18📚14📚16📚13📚17📚17Hand-picked from the consensus of the source lists: mostly entries cited by many lists, plus a few high-signal ones that fewer lists noticed. Each section links to its full catalog page.
The System One model family itself. Many members serve /v1/systemone, so the same client code works across them. docs/open-models.md maps the landscape and its caveats.
Hosted
/v1/systemone schema. See the Quick Start./decisions/v1/systemone and has a free tier. Reportedly #1 on the Jev Decision Index.Open and local
★29,237 · 📚38 — Convai's Apache-2.0 non-autoregressive encoder (421M/322M, 100+ languages, ~33 ms). Runtimes: laya-mlx ★6,654 · 📚23 (7–14 ms on M3 Max) and receptron/laya ★661 · 📚20 (Node/ONNX).★8,057 · 📚57 — A trainable family of System One models on Qwen3.5/3.8 with a pointer head and a /v1/systemone server.★4,615 · 📚63 — "Semantic ifs" from open models on a 3090, using shared-prefix logit readout.★2,453 · 📚60 · vinnylarouge/jevlike ★1,335 · 📚57 · Mapika/decider ★993 · 📚45 · bespokelabsai/nimble ★1,981 · 📚21 · wfzyx/von ★794 · 📚48 — Trained replicas and recipes (0.4B–35B), several with published limits.★986 · 📚32 · featherless-ai/simple-jev ★575 · 📚40 · razorback16/openjev ★548 · 📚49 · ekzhang/openjev-sglang ★335 · 📚48 · githubnext/localjev ★803 · 📚16 — Turn any open LLM into a decision endpoint, with no training. AnyJev adds option-order correction.★1,032 · 📚24 — "Ollama for decision models": it runs open System One models locally. It serves Laya, Decider, NLI and GLiClass behind a /v1/systemone-compatible API.System One decision models landed inside mainstream frameworks within days of Jev's launch, usually as a new evaluate / decide / classify model type alongside generate. Most of these integrations target the shared /v1/systemone shape.
★27,056 · 📚7 — The AI SDK's decision-model provider maps Choice, Score and Boolean onto experimental_evaluate.★5,426 · 📚28 — Vercel's agent framework. A decision model (Jev by default) powers its evaluate path and its tool-approval policies.★3,234 · 📚9 · vercel-labs/ai-cli ★817 · 📚23 — fx has a decision-model permission reviewer, reported at p95 18× faster than GPT Luna. ai-cli adds an evaluate command.★20,296 · 📚11 — A decision-model backend: output_type fields become typed questions (bool → Noul, Literal → Choice, IntEnum → Score).★59,945 · 📚9 — A decision-model complexity router, a compaction guardrail, and a pass-through proxy.★30,374 · 📚14 — A decision-model provider for tool and argument selection.★9,363 · 📚5 · BoundaryML/feelings ★22 · 📚10 — Maps the type system onto the primitives. feelings is a typed .feels() "AI if-statement".★1,240 · 📚7 — Records decision-model calls (Choice/Score/Noul) as OpenTelemetry spans. See also the Langfuse integration.★27,561 · 📚17 — A computer-use platform with a bounded-action recipe for decision models.★5,109 · 📚12 — A decision-model guardrail for jailbreaks, harmful content and secret leakage at the gateway.★2,445 · 📚16 — Nine search notebooks that use a decision model for ranking, filtering, routing and stopping.★8,206 · 📚7 — Uses a Noul to check whether a cached answer serves a new request.★41 · 📚20 · crmne/ruby_llm ★4,423 · 📚5 · laravel/ai ★1,203 · 📚6 — Java/Spring, Ruby and Laravel support.★143 · 📚23 — Makes decide a first-class inference type next to generate and stream, across 40 providers.★250,326 · 📚5 — A decision-model provider, skill routing, and a published (negative) compaction scorecard.📚3 — A production link-safety gate on typesafe-ai/jev.decide()), DeepEval, MotherDuck prompt_jev(), OpenRouter cookbook: gate tool calls, Ollama 0.35 decision models · → catalog/sdk.mdEvery major language had a community SDK within a week. Check the last commit date before you depend on one.
★6 · 📚30 — supports several gateways.★5 · 📚29★39 · 📚18 — SDK and CLI.★9 · 📚24★13 · 📚33 — async and blocking clients, observable retries.★6 · 📚31★1 · 📚23★0 · 📚6★51 · 📚32★7 · 📚35★18 · 📚30★16 · 📚24 — probabilistic control flow for Rails.★34 · 📚43 — pattern-match decisions inside OTP.★5 · 📚36★12 · 📚34 · Hawxy/TypeSafeAI.Net ★4 · 📚24★10 · 📚23★5 · 📚29★4 · 📚25★16 · 📚22★54 · 📚20 — maps Apple Foundation Models @Generable types onto the primitives.★7 · 📚20★96 · 📚48★2 · 📚30★39 · 📚8★20 · 📚18 — pandas and Polars.★6 · 📚27★8 · 📚23 — semantic rules inside Zod schemas.★3 · 📚22★25 · 📚39★13 · 📚32 — a CLI and stdio MCP server.★75 · 📚37 — Unix-pipeline and CI decisions with exit codes.★10 · 📚11 — jq meets decision models.★469 · 📚66 — The most-cited MCP server. Its typed tools (verify, screen for injection, find, rerank, classify, decide) follow a fail-closed contract. Install with claude mcp add jev -- npx -y @jkudish/jev-mcp.★337 · 📚61 — Formerly typesafe-mcp. A single connector for Jev, d1, CLM and Laya.★25 · 📚42 · burnigtm/jev-mcp ★61 · 📚29 · BYK/jev-mcp ★3 · 📚20 · PyModel/jev-judge-mcp ★69 · 📚17 · Brainwires/jevwire ★21 · 📚28 — Alternatives. BYK's server is eval-first, and jevwire adds an escalate-only Claude Code plugin.★47 · 📚18 — Classify first, read selectively. Batch text classification for agents.★18 · 📚17 — Hands off a whole browser task in one MCP call.★21 · 📚21 — An MCP proxy that screens both tool calls and tool results.jev-mcp repos share a name, so check the owner. → catalog/mcp.md★145 · 📚39 — The best community skill for writing decision-model programs: pick the primitive your code branches on, split two-property questions, and batch everything that shares a state.★932 · 📚42 — The most complete decision-model-in-the-agent-loop pack: routing, memory, compaction, skill selection, and computer and browser use (Hermes, Claude Code, Codex).★554 · source list — Five skills, a jev-decide CLI, 108 agent-supervision scenarios, and honest evals, including a negative agent result.★125 · 📚49 · ShivamPansuriya/jev-skill-gate ★7 · 📚24 · GodsBoy/jev-agent-skill-router ★23 · 📚33 · EliaAlberti/jev-rules ★62 · 📚27 — Show the agent only the skills and rules the current turn needs, and allow "none". jev-skill-gate cuts the skill manifest by ~75%.★33 · 📚34 — Hands agent steps that need no text output to a System One model (p50 ~230 ms). Its compaction and agreement evals are published.★36 · 📚30 · samtay32/jev-system-architect ★2 · 📚23 — Audit a codebase for fuzzy judgments that could become typed decision points.★2 · source list · Pleo2/awesome-jev-agent-skills ★0 · source list — Ready-made proceed/ask/stop gates, triage, diff review and QA-evidence skills.The largest cluster in the ecosystem. It covers Claude Code, Codex, Pi, Hermes, OpenCode and Cursor.
Context compaction and pruning. This is contested, so read the evidence first.
★7,250 · 📚71 — The most-starred project in the category. It replaces Claude Code's compaction summary with keep/drop decisions per tool call, and kept content stays verbatim. It began as a viral demo.★153 · 📚35 — Trims long Bash output before the model sees it.★99 · 📚39 — A calibrated context sieve with stub/recall pointers. It falls back to an LLM-backed adapter when the decision model is unreachable.★189 · 📚15 — Asks "is this a safe moment to compact?"★13 · 📚27 · leonaaardob/fast-dev-compaction ★9 · 📚26 · compozy/yoshi ★27 · 📚30 — Ports to Pi and Codex, and a pruning proxy.Model and effort routing. Route at session boundaries and run in shadow mode first.
★505 · 📚60 — Routes Claude Code and Codex to the cheapest capable model.★277 · 📚54 — Per-turn model, reasoning-effort and speed routing for Codex (~60% savings on a 237-turn backtest).★309 · 📚42 — One candidate set covering models, tools, subagents and skills.★16 · 📚21 · adarshmishra07/jcm-router ★8 · 📚27 · xinyao27/jevonian ★15 · 📚26 — Cache-aware routing that leaves the cached main chat alone.★293 · 📚15 · nidhi-singh02/agent-router ★99 · 📚25 · prismhq/jev-router ★14 · 📚29 — Mid-run effort selection, agent/CLI selection, and a router on LiteLLM.Guardrails and tool-call gates. Put deterministic rules first, and choose fail-open or fail-closed explicitly.
★48 · 📚38 — Deny/ask/allow auto mode for any coding agent. It uses session context, flags injection in tool results, and tracks taint. Smoke test.★153 · 📚50 — Steers instead of interrupting. Rule breaks fell from 6 to 0 over 150 paired runs, and a replay on 17k calls held 42.★618 · 📚61 — A supervisor for Codex and OpenCode workers, with risk-tiered thresholds.★30 · 📚38 — Auto-approves Bash, write and edit calls, and fails closed.★150 · 📚54 · Nyarlathoteppppp/pi-heed ★10 · 📚27 · harshwasan/jev-sentinel ★12 · 📚19 — Pi extensions that judge whether an action is destructive, exfiltrates data, stays in scope, or matches what you asked for.★5 · 📚19 · jesset/pi-verdict ★10 · 📚14 — Rules first, then a decision model for the grey zone. jev-engineering has a rerunnable 300-call injection test.★19 · 📚27 — Hermes command approvals: 8.7× faster, with 4.4× fewer prompts across 153 real commands.★32 · 📚27 · hemanth/pkg-gate ★0 · 📚16 — Screen codebases and npm lifecycle scripts before you run them.Done-verification, review and semantic lint
★646 · 📚74 — Staged code review (risk → file profile → evidence → severity → reviewer routing) with a local dashboard.★19 · 📚41 · qkal/Canny ★106 · 📚32 · noplan-inc/limpet ★5 · 📚24 — Stop-hook gates that block an unverified "done". Canny's rule: "deterministic hooks decide, the model advises".★469 · 📚19 — Enforces AGENTS.md rules with one Noul per rule.★12 · 📚39 — A pre-commit check that the message matches the diff; it blocks only on leaked credentials.★230 · 📚48 · supercorp-ai/supercov ★141 · 📚39 · lakeday-org/perch ★316 · 📚29 · mizchi/jev-lint ★113 · 📚23 · doeixd/jev-pref ★10 · 📚32 — Continuous quality review, coverage, and linting for semantic rules that a parser can't check.★2 · 📚19 · yamadashy/jev-labeler-action ★0 · 📚5 — GitHub Actions for submission review and issue/PR labels.The shared recipe: build a numbered list of what's on the page or screen (accessibility tree, DOM or OCR), let the decision model pick the operation and target, let code execute, and call an LLM only when text must be typed. Most decision models are text-only and can't see pixels.
★21,563 · 📚76 — The most-cited project in the entire ecosystem. It picks the operation and element in one request and booked a Zürich→London flight search in 7.1 s for $0.0039. The launch demo has been retested: 1/20 on complex tasks.★1,098 · 📚46 — macOS computer use: OCR the screen, then a decision model picks the action, for ~$0.0002 per step.★426 · 📚59 — Android automation; an Uber booking took 9 actions and ~21 s.★374 · 📚49 — Voice control at ~300 ms per spoken word. URLs are copied from your words, never generated.★297 · 📚49 · Ying-Kai-Liao/jev-browser ★93 · 📚43 · wy-coliney/jev-browser-use ★721 · 📚45 — In these, the LLM plans and the decision model decides. Ying-Kai-Liao's version solved 40/42 live tasks using ~70× fewer tokens.★613 · 📚28 · lahfir/agent-desktop ★1,733 · 📚16 · savka777/jev-use ★111 · 📚22 — Desktop control through accessibility trees, with local policy gates on sensitive clicks.★343 · 📚49 · realZachi/typesafe-adblock ★87 · 📚46 · anishfn/shapeshift ★755 · 📚23 — Browser extensions and UI: clutter removal, "is this DOM element an ad?", and a text box that morphs into the UI you mean.★495 · 📚64 — Web search where a decision model does source selection, query understanding and relevance ranking. It returns links, not generated answers.★386 · 📚59 — Ask Postgres tables questions in plain language. Its batching study found 20 rows per request gave 100%, and 80 rows gave 77–94%.★14 · 📚43 · colliber/duckdb-jev ★27 · 📚19 · mattn/sqlite3-jev ★3 · 📚15 · giuliosmall/pg_typesafe ★87 · 📚34 — Semantic SQL predicates. Every row leaves your machine, so pre-filter with cheap predicates first.★153 · 📚52 — Graph navigation: a Choice picks the next relationship and a Noul asks "goal reached?", with beam search in the app.★1,902 · 📚23 · keltokhy/jgrep ★131 · 📚32 · uehaj/sys1grep ★144 · 📚36 · sufianetaouil/every ★7 · 📚32 · ellipsis-dev/blink ★92 · 📚45 — "grep by meaning". These send code to the API, so check what leaves your machine.★487 · 📚31 — Fast document classification and packet splitting, from LlamaIndex's founder.★92 · 📚45 · RenaGao/jev-dataops ★60 · 📚18 — Dataset sifting (Rust, Parquet/JSONL) and streaming data selection with LoRA training.★8 · 📚33 · hev/reranker ★14 · 📚25 · hotchpotch/jev-reranker ★36 · 📚21 — Rerankers. Fuse with BM25 or embeddings rather than reranking with the decision model alone.★24 · 📚24 · keltokhy/jlink ★6 · 📚21 · keltokhy/jselect ★3 · 📚16 — Sort by meaning (pairwise comparisons plus Bradley-Terry), record linkage, and budgeted evidence selection.★16 · 📚33 · hyperspaceai/jevcache ★76 · 📚17 — OpenTelemetry log triage, and a decision cache for deterministic CI replay.★10 · 📚41 — Stop guessing thresholds. It fits per-question thresholds on your labels, verifies them on held-out data, and fails CI when a model update breaks them. Cited in more "best practice" sections than any other tool.★300 · 📚42 — Builds calibrated "AI functions" from human feedback, using active labelling plus GEPA to optimize the questions.★21 · 📚33 · jmanhype/jev-dspy-lab ★10 · 📚22 — Calibration, selective risk, and record/replay in DSPy.★98 · 📚26 — Agent evals and guardrails as typed decisions (it also runs locally with Kev or Laya): eight judgments per trace for ~$0.00006.★31 · 📚17 · nikkoxgonzales/jev-certify ★1 · 📚11 · sathariels/jevcheck ★2 · 📚7 · vcjdeboer/jev-reliability 📚4 — Criteria tuning, conformal routing bounds, behavioural contract tests that refuse jev-latest, and repeatability/phrasing preflights.★2 · 📚17 · ariel-frischer/jevkit ★3 · 📚22 · simota/tenbin ★4 · 📚18 — Lint your questions offline, before you pay for calls.★15 · 📚31 — Confidence gates, shadow mode, recipes and evals (on a row-filter job, Claude CLI took 48.9 s and the decision model 1.3 s).★0 · 📚6 — A conformance suite for /v1/systemone-compatible servers. Use it to check whether a Jev-like model really is a drop-in.The recipe for games and control: feed structured state (RAM → JSON, legal-move lists), never pixels. Let code own the physics, the route and the arithmetic, and let the decision model pick at branches. Keep a fast deterministic reflex layer that can veto it.
★422 · 📚64 — Super Mario Bros. from structured emulator RAM. It is the canonical "state as data" example.★228 · 📚66 — A MuJoCo drone. The decision model advises at 2.5 Hz while a 50 Hz reflex layer holds a veto.★9 · 📚38 — Code owns the route and the arithmetic, the model picks at branches in ~100 ms, and calibration is measured with Brier scores.★397 · 📚38 — 22 latency-focused demo apps by Nader Dabit.★27 · 📚45 · sorrycc/typesafe-snake ★23 · 📚32 · standardagents/jevpilot ★204 · 📚32 · rmalde/minecraft-agent ★568 · 📚18 — StarCraft, Snake (legal moves generated in code), a driving autopilot, and Minecraft (an LLM plans, the decision model acts).★249 · 📚30 · Dimweaker/jev-libero ★78 · 📚23 · TarunTomar122/jev-askable-arm ★11 · 📚27 · openroboto-ai/jev-robot-control ★53 · 📚13 — Robot arms that pick primitives from menus, never torques. OpenRoboto compares Jev with GPT-6 Astra on cost and time.★68 · 📚56 — Home Assistant: ask your house a question and get a number back, as sensors and automation actions.★2,708 · 📚71 — One trade decision per Monad block (~81 ms). It's educational: no list reports sustained trading profit, and bots should be dry-run by default.★483 · 📚41 — IRS form pages: 100% strict on 261 forms at ~$0.001 per page, replacing production Sonnet.★92 · 📚27 · parth-kp/jev-mail-classifier ★19 · 📚19 — Gmail triage: 1,000 emails in ~1 minute for ~3 cents.★7,181 · 📚37 — A read-only chat copilot for phones. An LLM drafts replies, a decision model ranks them, and you choose whether to send.★130 · 📚42 · AkashPriyadarshii/jev-seo ★88 · 📚32 · usenotra/notra ★225 · 📚19 — Social research, SEO/GEO audits, and a production GEO platform.★47 · 📚36 · backmeupplz/jev_antispam_bot ★13 · 📚20 — Discord and Telegram moderation.★12,344 · 📚29 — An AI trading OS that uses a decision model as an evidence and risk gate. It blocks entries but lets exits and protective orders through.★37 · 📚22 · youkiti/tiab-review-plugin — Systematic-review screening (95% recall on 16,645 records) and extraction.★244 · 📚33 · ChetasLua/jevmeter ★102 · 📚45 — Fun ones: kill, fix or ship your startup idea, and a live decision meter on any video.Summarized with numbers in docs/evidence.md.
★189 · 📚40 — JevBench: 534 frozen, half-sealed decisions and a four-axis composite across 40+ systems. See also the Benchmark Heaven leaderboard.📚12 — Jev vs six LLMs on human-labelled sets, with the data published.★0 · 📚11 — 123,805 pre-registered requests. ECE was 0.075 in-domain, but calibration fails OOD and authority injection works.★7 · 📚32 — The canonical question-atomization lesson: 62.6% with one question vs 95.0% atomized. Haiku 4.5 wins the direct question.★5 · 📚12 · bitnovus/jev-spam-eval ★2 · 📚34 — Classical baselines win in-domain, but Jev wins under drift.★9 · 📚39 · zhuyansen/jev-search-rerank-eval ★9 · 📚31 — Reranking: Jev ties Cohere, but doesn't beat embeddings alone, so fuse them.★0 · 📚15 · KantaHayashiAI/jev-does-not-play-dice ★3 · 📚17 · scienthoon/jev-ood-calibration ★6 · 📚25 · jourdanlabs/assay-001 ★0 · 📚18 — Calibration audits covering abstention, known probabilities, OOD, and a pre-registered test.★3 · 📚31 · zkousama/jagged 📚4 — Prompt injection and vulnerable-code detection.★86 · 📚18 — WebMCP: Jev + Mercury solved 49/49 tasks, bare Jev 25/49.★26 · 📚27 · RINNECODER/jev-behavior-study ★3 · 📚23 · FirasSX914/Janus ★2 · 📚25 — Where Jev holds up vs breaks, order and lost-in-the-middle effects, and when routing is worth it.📚21: one judge call vs dimension scores.📚19📚14~35 papers on Jev and other System One models appeared within two weeks, all unreplicated preprints. The full annotated list is in docs/papers.md.
📚18 — An ecosystem survey of 2,170 GitHub projects.📚7 · JEV vs. LLMs as Rubric Judges 📚5 — Jev comes within ~3 pts of GPT-6 at ~0.36% of the cost, but LLMs repeat its confident errors, which caps cascade gains.📚5 · JevOut 📚4 · JevAdvBench 📚5 · Decision Hijacking 📚6 — Robustness: option names, natural context and injection.📚4 · Jev-Mobile 📚4 · JEV-Star · Jev-Mem 📚6 — Planner/executor agents that make 66–73% fewer strong-model calls.📚4 · Just Ask Jev 📚7 — Annotation, large-scale coding and alignment-failure detection.Analysis
📚11 and 6 clones of Jev in 2 days.📚17 — Often named as the best long-form introduction.📚8 · TrueFoundry: What actually shipped · Silverthread Labs: What the 193× benchmark doesn't measure · ContextOS: production engineering review — Critiques.News
Community discussion
★124 · 📚10 — From one smart if-statement to a coding agent that reaches for a decision model on its own.★14 · 📚15 — A free hands-on course: 13 agent use cases, a fast brain and a slow brain, one OpenRouter key.★34 · 📚27 · datawhalechina/jev-cookbook ★46 · 📚14 · paramjeetn/jev-cookbook ★7 · 📚11 — Tested recipes on OpenRouter, 18 Chinese notebook recipes with evals and local fine-tuning, and 120+ use cases.★33 · 📚14 · Foadsf/jev-for-engineers ★5 · 📚23 · PromptEngineer48/langchain-jev-tutorial — An illustrated primitives guide, engineering examples, and a LangChain support-ops agent.📚9📚14: builds with reported cost and speed.📚4How it was built
owner/repo, and links to the source lists themselves are dropped.📚 counts how many distinct lists cite it.Caveats
Updating
scripts/refresh.sh # re-pull all sources → stars/commits → catalog → sources.md → README numbers
To add a source, append its GitHub URL to sources.md and refresh. The hand-written synthesis in docs/ and templates/README.md.tmpl should then be reviewed against git diff catalog/.
Repository layout
README.md ← generated from templates/README.md.tmpl (live ★/📚 numbers)
sources.md ← the 106 sources: stars, pull date, commit
docs/ ← synthesis: what-is-jev, patterns, evidence, open-models, papers, source-lists
docs/source-notes/ ← one review note per source list
catalog/ ← the complete categorized union (18,330 entries)
data/ ← catalog.json/csv, sources.json, verified.tsv, source_profiles.json
scripts/ ← pull_sources, extract_links, build_catalog, verify_github, render_*, refresh.sh
Python
94.3%
Shell
5.7%