A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
Python
2,061
678 commits
updated Oct 1, 2026
A curated awesome list of public projects and practices built on Jev, TypeSafe AI's System One model for typed decisions.
This README is the homepage aggregate of the current category files, so the latest accepted entries are visible here without drilling into subpages.
Jev is not a chat model. It takes unstructured state plus a typed question and returns a typed decision — a choice, a score, or a boolean, each with a confidence. That makes it a drop-in decision layer for software: classification, routing, rubric scoring, verification, and agent guardrails. This list tracks who is actually building with it, and which patterns transfer across industries.
The repository treats all categories equally — each entry lives in exactly one category, chosen by its direct Jev application domain. A dedicated Related Practices / Discussions category captures credible public practice signals — X threads, Reddit discussions, and interviews — that describe real Jev usage even when no strong standalone case page exists yet.
[!WARNING] A listing is not an endorsement. This project applies inclusion rules only — public, citable, genuinely uses Jev for a typed decision, one-sentence summary. It does not review code quality, security, maturity, or whether a project runs at all.
Treat same-day bulk submissions with particular care. Several repositories published together by one author, sharing a scaffold and a thin commit history, can satisfy every inclusion rule and still be unproven. Volume is not evidence of quality. See Curation is not endorsement for a checklist to run before adopting anything here.
Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. This list answers two practical questions quickly:
This is not a comprehensive database. It is a high-signal, fast-scanning field guide.
An entry should meet all of the following:
Jev/jev, cites TypeSafe AI's System One models, or shows a typed-decision loop (typed question → typed answer with confidence → accept/reject/escalate).We do not include:
Inclusion means one thing: the entry satisfies the inclusion rules above. It is not a quality review, a security audit, or a recommendation. We do not verify that a project compiles, that its tests pass, that its published numbers reproduce, or that its license permits your use.
This matters most for projects that arrive in bulk. When one author releases several repositories on the same day, they commonly share a single scaffold — the same AGENTS.md, CLAUDE.md, STATE.md, and CHANGELOG.md — land in one or two commits each, and may ship considerably more prose than code. Such projects can be entirely legitimate; they are simply unproven. Treat them as leads, not as validated tools.
Before adopting an entry, check it yourself:
| Check | Why it matters |
|---|---|
| Does the code actually call the Jev API? | An entry can read well on a README alone. Look for a real request carrying typed questions, and a parsed answer coming back. |
| Is there a runnable check? | A test, an example with expected output, or a public demo. No check means no evidence that it works. |
| Do the numbers have a source? | Any accuracy, latency, cost, or volume figure should be traceable to the linked page. We strip claims we cannot verify, but the project page itself may still carry them. |
| How much of the repository is code? | Some projects are mostly prompt documents. That can be legitimate — just know which one you are getting. |
| Is there a license? | A few entries have none, which limits reuse and redistribution. |
Found something wrong? Open an issue or a pull request — removal is as valid a contribution as addition. Rules for AI-assisted work, project depth, and submission rate live in CONTRIBUTING.md.
Each entry lives in exactly one category. When a project could fit multiple categories, we choose the one closest to its direct application domain.
Optional tags on an entry name the coding agent it targets and the kind of integration it is. Most entries carry none — they are added only when the source itself supports the classification.
Source file: categories/classification-routing.md
NOTRA_JEV_CLASSIFIERS flag routes brand-visibility classifiers off an LLM and onto Jev Boolean decisions at a 0.5 threshold, targeting 300 ms p50.Choice (invoice or general) and routes each inbound email to the matching handler./v1/systemone endpoint to new-api so typed decisions sit behind the same gateway as chat models.Choice and Noul questions, sending low-confidence answers to human review.Choice picks plain code, Jev or a reasoning LLM behind a Noul gate for non-tasks, code vetoes Jev when the idea needs images, and low confidence returns "not sure"; closed source, free page and API.Choice over the installed skill roster plus Boolean-style gates on whether any skill is needed, suggests a skill only when the gate and the per-candidate fit both clear 0.30, and defaults to a shadow mode that logs the decision without injecting it.Choice over ten kinds of post plus three Noul questions (paid ad, clickbait, emotional pressure) about each, counting an ad from 0.7, or from 0.4 when the kind is also ad, and clickbait and pressure from 0.5, then draws the monthly mix on a shareable card that links the highest-scoring posts for a manual check.Choice questions to classify CSV columns into a 13-code type vocabulary and each dataset into one of six scenes, then executes every write locally; measured Jev at 6.6–12.7× an LLM's token cost on this task because the output is already one character while per-question criteria repeat.jevonian/auto from session state, quota health, candidate capabilities, and cache-switch penalties, after deterministic code has filtered candidates and while pinned models, explicit jevonian/<route> requests, and routing.mode: "off" skip Jev entirely; minConfidence marks a low-confidence route in the ledger rather than accepting it, and the ledger records the serving model and why.Choice question per tab against user-editable group criteria — with unclassified tabs falling to a fixed fallback bucket, manual groups untouched, and the pre-grouping tab order restored on ungroup; an optional LLM engine with self-invented group names is included for comparison runs.Noul judgments against per-platform, user-defined topic and expression labels to annotate Weibo, Threads and X posts directly in a Chrome extension.Noul per remaining (source, destination) pair plus a guard Noul per incoming column, mapping 10 of 10 columns of a 23-column export at 253 questions in one call, 915 ms, $0.0012, unmapped fields left visible above a 0.75 threshold rather than guessed.Noul actionable, Score severity, Choice owning team, Choice duplicate-of), pages on P(SEV1)+P(SEV2) ≥ 0.80 and drops only below 0.20 when actionability agrees, sends the band between to a human who has 15 minutes to ack before it pages anyway, links duplicates into a cycle-broken incident graph so two alerts can never silence each other, and falls back to configured severity on any error or timeout; a published 300-alert run measured p50 418 ms, p95 1477 ms and $0.04 per 1,000 alerts, with the slowest call landing 151 ms short of the 2 s timeout.Choice and Noul questions about issue category and reproduction steps in opt-in live mode, then applies confidence and reproduction gates to propose a queue or review fallback without assigning the issue, with synthetic offline fixtures and policy tests.Choice, Score and Boolean questions to sort the peer's message into intent (10 options), an emotion distribution (9 options), urgency (0-3 Score) and a suggested reply posture (11 options), then shows exactly three toasts and takes no other action - no generated reply, no input injection, no screenshot or OCR, with a local alias-to-relation table passed as state so the same sentence is judged differently for a partner than for a colleague.Choice over member descriptions, fills blank task fields with Choice and Score questions, and routes automation workflows on a Choice/Score/Noul condition node, applying answers only at 0.6 confidence or above and otherwise leaving the task unassigned or taking the Else branch.Choice per user rule plus a Noul on whether the page is a payment, login or banking screen, stepping in with a pop-up only when a rule's probability clears its threshold and never on sensitive pages; on 119 trial pages with Kev, the short-video, feed, livestream and video rules had precision 1.00.Choice over the readings of a Chinese polyphonic character while the model stays fixed and only the harness around it is optimised, ending at half the errors for a seventh of the cost.Choice over /effort levels (low / medium / high / max / unclear) plus a Noul on whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set.Noul on the target plus Score rubrics about each row's text, turns every option's probability into a column next to the row's numeric fields, and lets a tabular foundation model such as TabPFN learn from the labeled rows in context, reaching 0.745 AUC at 256 labels on Kickstarter funding against 0.682 for Jev alone with calibration.Choice questions for kind, project and next action plus a Noul for duplicates, routing to a project only at 0.45 or above and marking a duplicate only at 0.70 with a specific matching item, while the text itself is stored verbatim in append-only local files.Source file: categories/adaptive-realtime-ui.md
Noul question, batched 20 at a time, answers whether it reveals a concrete plot event of the video being watched or of another title the user protects, with the extension owning the 0.85 / 0.7 / 0.5 threshold and keeping the comment covered when the check fails.Noul per menu item against the user's plain-language request, and presses the top match when it clears a probability threshold, falling back to a ranked list otherwise and never auto-running destructive items.Source file: categories/verification-guardrails.md
risk-check provider where Jev scores agent counterparties as typed Noul/Choice/Score questions into a code-controlled 0-100 score, issuing an ES256-signed attestation per verdict; 540-call scale run (99.76% at threshold 65-75, 0 false positives) and a 5-iteration 1,500-case adversarial red-team loop (100% adversarial accuracy) with ~$0.00005/decision at p50 ~400ms.Choice over the cross-turn coding-agent trajectory decides continue, pause, or escalate, with low-confidence verdicts routed to a human while deterministic code keeps dangerous-command blocking, budget caps, and an Ed25519-signed receipt chain.Noul checks about source and build files, escalates suspicious chunks for a second pass, and returns implicated files and lines before execution.typesafe_permission_reviewer builtin so the agent's permission decisions run through Jev rather than an LLM call.Boolean questions per paragraph (stacked hedges, restating closers, not-X-but-Y turns, naked cost figures) at a 0.7 threshold; CLI, pre-commit hook, GitHub Action and Claude Code skill; measured 182 ms median and 1 of 54 clean paragraphs flagged against 37 for Haiku 4.5.jev-pref.json rules that Jev checks against each diff hunk, staged file set, or pull request, returning fix_now or advisory findings to the coding agent and a nonzero exit code on blocking ones.Noul risk checks plus a risk Score, maps the answers in code to allow, ask or block by the user's risk setting (protected paths always ask), and also uses Jev to pick the model tier per prompt and to send back "done" claims that ran no verification, at about 400 ms per decision.Noul checks per tool call (beyond scope, destructive) and admits each judgment as Evidence that can degrade Trust, Delegation, and Authorization until a human override repairs the relation, so later calls inherit the history; includes a stateless-vs-stateful comparison with a scenario adversarial to persistence.Noul review question for each translated subtitle line to flag omissions, changed meaning, names, numbers, negation, or other defects for human review before export.Choice with an explicit Unresolved option to distinguish a mentioned arrival time from an agreed meeting time, showing the returned probabilities and prompting for missing agreement before treating a time as settled.Noul+Choice call per fact while deterministic code keeps the confidence gate; 90-sample calibration reports 0.922 gated accuracy and 0/15 injection flips.typesafe-ai/jev in malicious-link-check.ts before a short link is created, so the URL is gated by a typed verdict rather than a blocklist.Noul, Choice and Score questions about one function, file outline, candidate copy pair or test at a time, turns answers at 0.80 into review or consider findings with file and line, fails the build on review, and keeps undecided files as uncertain instead of clearing them.Choice risk class plus Noul irreversibility, task-match, and injection checks gate every non-read-only tool call (deny confident high-risk, ask ambiguous, delegate the rest), Noul scope and reversibility questions auto-approve clearly-granted calls behind argument-evidence gating, and a per-message Score re-ranks what a referenced session keeps instead of oldest-first dropping — fail-closed to stock behavior, shadow mode with a /jev-stats command, 64 tests.jev-guardrails mod that screens prompts and turns against jailbreaks, harm, and policy breaches via TypeSafe System One.Score lenses and one Noul — and blocks a diff, commit message or doc set that falls below the resulting quality score, from the terminal, a pre-commit hook or a GitHub Action.Noul judgments and local thresholds to flag team-rule violations as Claude Code and Codex edit, helping agents fix them before code review with configurable rule packs and repository-specific rules.Noul on whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable; 6 of 6 secrets and 0 of 6 benign strings were blocked in its published calibration.Noul for whether it has a bug, a Choice for which kind and which line, and a Score for severity, plus language-filtered CWE Noul checks and custom rules written as sentences at repository, file or method level, failing CI on any answer over its floor.Noul for each piece of code a rule applies to and reporting it above the rule's threshold; its two shipped rules were right on 12 of 12 sampled findings in three open-source projects.Score plus Noul checks for answer evidence and prompt injection, then withholds any draft whose claims fail a batched per-claim Noul or cite numbers absent from the sources; on its replayed 62-question golden set it answered 0 of 22 unanswerable questions, against 4 of 22 for naive top-5 RAG.Source file: categories/scoring-ranking.md
Noul to score already-loaded X posts against a reader's goal and profile, then pairwise Choice judgments to rank eligible matches and surface up to three for review.Score axes inside a single systemOne request and turns them into a 0-100 Slop Score in ordinary TypeScript.Noul properties per source file so the agent knows what to fix first.Noul checks to rank candidates, apply evidence thresholds, or extract source text while keeping those decisions independent.Noul judgments to find code, docs, logs, and text satisfying natural-language conditions, with a configurable probability threshold and ranked file results linked to source lines.Noul per line on whether it answers the query (16 lines per request, sharing the page text as state), highlighting lines at or above 0.55 page by page and ranking them by probability.Choice to locate source lines and Noul to check evidence presence across document regions in parallel, then returns original passages to a GPT research agent through Pi-Serini with 20/40/60-document batch limits.Noul relevance question per chunk across three parallel requests, and returns the accepted ones as excerpts through MCP; the cutoff and excerpt selection live in code.Score per item per weighted dimension of a YAML spec in a single request and sums weight times score in code to order the items, with noul, choice and score commands that turn a threshold into exit code 10 for shell scripts and CI.Score over each experiment result to decide whether it clears the promotion bar, and a Choice over candidate plays to decide what to run next in SEO, content, and ads.jev-latest against api.typesafe.ai and treats the returned probability as relevance, because TypeSafe exposes no native rerank endpoint.Noul per page on whether the visitor would be glad to land there, a Choice for the single best answer, and a Noul on whether any page answers at all) and re-orders or drops hits in code, with the repo's own benchmark over the 109-page TypeSafe docs reporting Hit@1 of 83% against 41% for its keyword pass alone.jevr) that turns a natural-language query into grep-style path:start-end targets — a stateless local BM25 pass proposes candidates, Jev Noul membership questions score their 100/20-line windows (kept at 0.90 for code, 0.60 for docs), and one listwise Choice per lane orders the keepers — ships as a Claude Code skill and plugin, and placed 2nd of 90 models on the HAKARI-Bench NanoRTEB reranking leaderboard.priorityGap that a local uncertainty policy turns into the order a human should read the hunks in, while OpenAI explains the ones that surface.Source file: categories/agent-decisions.md
Choice (allow / ask / block) with Noul irreversibility and exfiltration checks before every tool call, sends ask verdicts to a human and fails closed on errors, and adds a Choice model router with a confidence fallback and a Noul "am I done?" gate, each measured against labeled fixtures.Choice at each step to select a concrete socai CLI operation and observed post or profile target on Instagram, TikTok, or LinkedIn, rejecting malformed or low-confidence decisions before execution.Choice (wake, not yet, unrelated) against the agent's own sleep note, and code skips the wakeup only when wake is below 0.2 while always waking on user messages, bare timers, a skip limit, errors, and timeouts; one run passed 21 of 21 hand-written scenarios, which the README calls a smoke test rather than a benchmark.chrome_act_toward_goal) with an 85%+ pruned DOM tree (Shadow DOM & iframe pierced), dispatching native CDP events (isTrusted: true) on active logged-in sessions without focus theft.Choice picks the operation and DOM element each step and two Noul checks (goal reached, stuck) veto a premature DONE or BLOCKED, with a small text model used only when text must be typed.Choice per step (operation and target control) over the front window's accessibility table, executes only validated high-confidence answers, and escalates to a full LLM agent on low confidence, no-effect actions, or unknown field values; author-reported 275–690 ms per decision.Choice picks the next tool from candidates rebuilt every step, a Score grades the call's risk, and a Noul decides whether it needs authorisation, while plain code acts on the answers so a high risk score forces human authorisation that no probability can override (7.7% of wall clock with the offline judge, 79% over the hosted API).Choice decisions select tools and targets, uncertain decisions escalate to an LLM, and a shared guarded kernel supports isolated Docker workspaces and paired LLM-only comparisons.Choice over grid-tile candidates at each level, and local probability and margin gates decide whether to descend or refuse, emitting only a raster point and bounding box and never clicking.Choice over candidate replies, and fills the draft box while sending stays manual; a Windows port does the same from offline OCR of the WeChat window.jev-decide script and ships a P0 case set taken from accessibility-tree snapshots of a calculator, a calendar, and the NetEase home screen.jev_decide tool lets the agent ask a Noul, Choice, or Score question about any state — urgency triage, intent routing, guardrail checks — and gate on the returned probability or confidence in code instead of trusting the chat model's guess.Choice picks the operation and its target element each step from candidates read off the DOM rather than the layout, so the same agent runs on Chrome and on engines that never draw a page (Moli, Lightpanda, Kitesurf), passing at least 90% of runs on each of fourteen browsers tested.Choice, a Score and a Noul on each prompt and after every tool batch to pick the next API call's reasoning effort (low/medium/high), sent as a per-message statement so the prompt cache never breaks.Noul on whether a shell command is strictly read-only, auto-approving at 0.95 and otherwise falling back to the normal permission prompt without ever denying, while a local hard-no list and injection filter keep risky commands from reaching Jev; 0 of 8 state-changing commands were approved in its published calibration.Choice questions for the next operation and its target element behind the same /v1/systemone API, completing 38.5% of 125 hand-picked real-website tasks graded by deterministic verifiers versus 16.7% for Jev 1.13 in the same agent.Source file: categories/data-labeling-curation.md
Choice, Score, or Boolean decisions, sends ambiguous and audit samples to a human, and uses accepted human labels to optimize the saved definition with GEPA.tail -f output, by asking Jev one Noul per line against a plain-English question and printing lines at or above a probability threshold, with a hand-labelled benchmark against Claude in the repository.Choice per span (a schema class or none) plus a Noul per sentence-level class, keeping answers above a per-class threshold and flagging close calls for review, with a published benchmark measuring 10–26× lower cost than LangExtract on Gemini 3.5 Flash but lower F1 (84.2 vs 88.5 on its bilingual jx-bench).Choice over every candidate window per entity type, verifies each nominee with a second Choice (the type, none, mixed or partial) and settles its boundary with a third, averaging 73.7 strict F1 across 12 Chinese and English NER benchmarks against 72.1 for direct extraction with Qwen3.8-27B.Source file: categories/evaluation-benchmarking.md
Choice questions about first-visit understanding, returning inspectable findings for the first change to make.state costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing instructions in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.Noul rather than Choice moves the result more than the gap between the two models, at a thirty-fourth of the cost, written up at jjd-lab.github.io.Choice over the six faces of a hidden fair die 400 times; Jev selects face 1 on all 400 trials with 82.9% mean reported probability and 19.0% accuracy, then tests whether stated probabilities survive in synthetic forecast documents, where a 30% shortage risk comes back as 5.3% via Choice and 26.7% via Noul; raw responses and analysis code on GitHub.Choice (same / fact_differs / action_differs / specificity_differs) decides which of an agent's approved answers changed meaning rather than wording, and on 109 before/after pairs whose ground truth is derived from what each config rule does to the answer, Jev catches all 19 real changes with 13 false alarms against 33 for a markers-then-embeddings-then-LLM stack and 19 for the LLM judge alone.jev-1.13-20260917 through OpenRouter's TypeSafe-compatible /systemone endpoint, reporting about 261 fixed input tokens per request, zero spread in the implied per-request cost across question counts, batched-vs-single answer differences comparable to repeat-request noise, and median input-token savings of 76–86% at eight questions.Choice/Score/Noul or any OpenAI-compatible backend (with a free rules fallback), gates low confidence at 0.7 (caught 3/3 misjudgments at 9% escalation, n=130), and publishes Chinese-scenario cost-accuracy numbers — 97.7% @ ¥0.105/1k decisions and 60.0% → 68.3% on a frozen 120-item human-labeled spam set at τ=0.10.Choice on 108 verified four-option questions about the Convex backend platform (no docs or tools in the prompt, each asked 3 times with shuffled options, random guessing 25%) alongside 14 LLMs, where jev-1.13 scores 84.6% at a 199 ms median and $0.0088 per full run against 98.0% at 2.12 s and $1.59 for the top model, with every answer, probability and raw request/response in a public explorer and the runner in get-convex/convex-evals.Noul on whether an answer misrepresents its PubMed abstract; Jev scored 92.9% against 92.4–95.1% for four fast LLMs at a 204 ms median and USD 0.03 per 1,000 checks, and letting Jev settle the 37% of items where it was at least 90% sure kept each LLM's accuracy with 37% fewer LLM calls.Choice(2) up to Choice(16).Source file: categories/calibration-research.md
choice, score, and noul questions with RLCD-trained calibrated probabilities in a single ~35 ms forward pass, published on PyPI and Hugging Face.POST /v1/systemone server answering typed Choice/Score/Noul questions with confidence from open weights, verified as an official-SDK drop-in with temperature-fit calibration (set3 n=1316, 0.83 overall)./v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.POST /v1/systemone server that reads typed Choice/Score/Noul answers from the logits of any MLX checkpoint or OpenAI-compatible endpoint with no training, works as a drop-in for the official SDK, and replays Jev's published answers on 256 public judgments (231 vs Jev's 238, McNemar p = 0.21).Choice (2-255 candidates), Boolean, and Score questions with dev-calibrated distributions in one forward pass and zero output tokens, trained from scratch on CPU with committed datasets, predictions, and ECE results (maze 0.016).Choice/Score/Noul interface on commodity zero-shot NLI models and makes the confidence honest with temperature scaling and conformal abstention, shipping a reproducible calibration eval (ECE 0.170 to 0.071, cross-validated) that runs offline with no API key.Choice, Noul, and Score questions from structured state and documents without a hosted call.Choice, Score, and Noul in a single forward pass and returns calibrated confidence meant to be thresholded, so cases it is unsure about escalate to a human instead of being guessed; MLX-first on Apple Silicon, with a System One endpoint and weights on Hugging Face and ModelScope.Choice over 13 options (0–9, ., -, END) per answer character given the expression and the digits so far, appending whatever it picks and showing each step's full probability distribution and confidence, with a Rerun that exposes call-to-call jitter on identical inputs, live on Cloudflare Workers.Choice, Score, and Noul questions on the same POST /v1/systemone wire format, calibrated with temperature scaling plus a split conformal abstain set with a coverage guarantee (ECE 0.01 to 0.03 on the public suites), runs on CPU or in the browser via ONNX, fits on your own labels in seconds, and its README says it loses to Laya on typed decisions (0.71 vs 0.77).Choice questions one at a time, with confidence driving lookahead when the top pick falls below 0.65, beam search across sentence directions, and a self-critique loop that rewrites sentences scoring below threshold; live at talktojev.com, paper at doi.org/10.5281/zenodo.22940945.Noul, Choice and Score on the same POST /v1/systemone wire format in one forward pass, scoring 0.791 top-1 on typed-decisions' 2,000 test decisions after training on its train split (Jev 1.13.0: 0.727), and 0.774 on 873 scienthoon support tickets it never trained on (Jev: 0.753).Choice and Noul questions behind a TypeSafe-compatible API, matching Jev across computer-use, gaming and tool-calling with up to 9x lower latency and reporting 87.6% on Terminal-Bench 2.1 as a fine-tuned verifier.Choice, boolean or rubric-score question, released with its data pipeline, training config and eval harness under Apache-2.0, and reporting 90.1% on its 324-example holdout against Jev's 93.2%.Choice, Noul and Score questions with a probability for every option from one forward pass, and run anywhere: a standalone server with Jev's /v1/systemone wire and a playground, a llama.cpp script for the GGUF builds, or AINode (open-source local AI platform).Noul yes/no questions on Jev's own /v1/systemone wire format, running CPU-only via llama.cpp at 54 ms short / 220 ms long on a laptop Core Ultra 7 255H against Jev's hosted 344/345 ms; Choice and Score return 422, and it scores 0.815 accuracy on 2,000 unseen policy yes/no questions against Jev's 0.927.Choice, Noul and Score questions via prefill-only inference on NeoHorse-1-4B, deployable with vLLM, SGLang or a native Python/CLI/HTTP runtime, scoring 77.70 across six text benchmark groups (highest among open-weight entries with complete results in its published comparison).choice, noul and score on the same /v1/systemone format at about 22 ms per decision, published with a panel that measures Jev itself at 0.828 accuracy and 0.053 ECE while stating it claims no statistical significance.choice, noul and score on /v1/systemone with per-checkpoint provenance and calibration plots released./score call per step and supports per-task fine-tuning, released under MIT with the terminal run recorded.Source file: categories/infra-sdks-integrations.md
typesafe-ai/jev) in its experimental evaluate path.evaluate command.Choice, Score and Noul questions on Huawei 910B hardware — 37–47 ms median for a four-question request, 33.8x–70.9x faster than the same request on a single container CPU thread, with an output-equivalent SDPA decision head that avoids torch_npu's CPU fallback on aten::_transformer_encoder_layer_fwd.Scores each retrieved passage and Choice/Noul selects the query engine, with nfcorpus nDCG@5 0.340→0.396 at about $0.0003/query.npx skills add typesafe-ai/skills) that teaches agents the Jev workflow.jev-browser), so Jev arrives as a first-class Cline capability.Choice, Score and Yes/No questions to the model about pasted text - ticket triage, moderation, review scoring - and returns a parsed answer with a confidence value in about 0.5 s per decision.jev() filters, jev_prob sorts, and jev_choice groups.Choice/Score/Noul question sets with 13 offline lint rules before any Jev call, then sends the canonical wire payload and prints parsed, confidence-bearing JSON answers to stdout using exit code 2 to reject a billed-but-useless request.Noul, Choice and Score answers into named decisions with enter/exit thresholds (hysteresis), nested decision trees settled in one call, a JSONL journal, replay of a threshold change over recorded answers with no inference, and Brier/reliability calibration, over TypeSafe direct, OpenRouter or Vercel AI Gateway.Choice, Score and Noul questions through a decision client separate from the text model, the Kernel picks a specialist subagent by Choice, and each retrieved RAG passage is withheld from the generating model when its prompt-injection Noul exceeds 0.70.if Hunch.likely?("fraudulent", given: order) reads like plain Ruby but branches on a typed Jev answer, with pick for Choice, rate for Score, and graded predicates from possibly? to almost_certainly?; a TypeScript port offers the same interface.decide as a first-class inference type alongside generate and stream.Choice, Boolean, and Score decisions on pinned open models across Torch, vLLM, MLX, llama.cpp, and WebGPU, with committed row-level benchmarks and checksums.Choice, Noul or Score answer becomes a typed branch under caller-supplied thresholds, anything below them takes an Uncertain case the compiler forces you to handle, and procedure routing skips the model call entirely when deterministic predicates leave one candidate.Noul, Score, and Choice decisions into deterministic threshold workflows that batch into a single systemOne call and return an ordered, explainable action set instead of side effects, with matched rules recording the actual value behind each action and a mock provider so policy tests run without an API key.Choice, Score and Noul decision, trains a head per decision site, and answers live with calibrated confidence, falling back to the upstream below its threshold and demoting itself on drift; measured 22 ms per decision and 90% of support-triage questions answered locally at a 0.99 agreement target.Choice, Score, and Boolean evaluations; recipes display answer probabilities separately from provider confidence and use deterministic local thresholds to pause uncertain routes for review, while every configuration exports as TypeScript.Noul, Choice, and Score service that shares one state prefill across isolated questions and, on its included 27-question Qwen3.5-9B/A100 fixture, reports 0.500 s versus 13.554 s for generated JSON.Noul, Choice and Score questions that reads Choice and Score answers back as the caller's own types, rejecting any label or level the question never offered, with retries, OpenRouter support, and eleven examples tested against an in-process fake of the API.POST /v1/systemone wire contract, each rule citing TypeSafe's docs, OpenAPI file or SDKs, and a suite that checks any Jev-compatible server against it (Choice probabilities keyed by option and summing to 1, Score equal to Σ i·p, 2–255 options, error shapes, answers that stay put when question ids or order change), finding 2 of the 8 most-starred open ports conformant, with a reference mock that breaks each rule on purpose and a proxy that fixes what can be fixed.Noul, Choice, and Score questions and returning named answers as pipeline-friendly properties.TypeSafeModel integration to map Pydantic schema fields into typed Jev System One questions with confidence scoring.#[JevNoul], #[JevChoice]), Workflow guards, and WebProfiler panels.pip install "jev-style[torch]" (or [mlx] on Apple silicon) serves an open 0.8B Qwen3.5 decision model behind a System One-compatible /v1/systemone API that answers Choice, Score, and Noul questions with calibrated probabilities in one pass over inputs up to 25,600 tokens (0.15–0.2 s per short request with MLX on an M1 Max, after the first call), and ships a Claude Code guard that turns four Noul checks and a risk Score into allow, ask, or deny through code-owned thresholds.grev 'is a vegan meal' menu.txt keeps the lines whose Jev Noul clears 0.5, sibling filters route by Choice and rank by Score, and the output is always your own input, never generated text.Retry-After delay when the service answers 429, refuses any request that would pass a per-key daily spend ceiling (USD 0.20 by default), and keeps an opt-in cache of answer probabilities under caller-chosen keys, so a repeated question over the same state is not paid for twice./v1/systemone endpoint, so an existing Jev client only has to point TYPESAFE_BASE_URL at the daemon, and reports its recommended model at 0.722 accuracy against Jev's 0.738 on typed decisions.WHERE jev(people, 'the name is European'), jev_prob and jev_choice each put one Jev judgment per row with calibrated probabilities./v1/systemone proxy and Pydantic AI client that records every Jev answer's probabilities and alerts, without labels, when jev-latest switches versions, a question's answers drift (chi-square-tested PSI), answers crowd a decision threshold, or estimated accuracy falls.Choice, Score and Noul next to parallel fan-out, a router and a guardrail in one codebase.Choice, Score and Noul questions onto POST /v1/systemone from a separate provider package, with a runnable ASP.NET ticket-triage sample that picks the owning team and escalates to a human at a normalized Score of 0.8, and 1,096 tests across net8.0 and net10.0 that run with no HTTP.Source file: categories/game-simulation.md
Choice judgments to answer lateral-thinking puzzle questions and assess proposed solutions, with application code requiring supported facts, a coherent explanation, and sufficient confidence before marking a puzzle solved.Choice between two names under land, water or air rules held in state, asking the champion against the next K challengers in a single request and discarding the speculative answers once the champion falls — 1,999 fights in about 16 s at roughly US$0.01.Choice over four directions with no heuristic fallback, gated by a user-set confidence threshold that pauses play for human review, with editable prompts and board rules, bring-your-own-key backends, and archive import/export.Score questions on who would care plus seven Noul moderation checks (0.5 keeps the text out of the public feed, 0.85 blocks it), batched Choice questions then return each persona's reaction in waves of 600, 1,500 and 3,000, and code sends the text to the next wave only while glad reactions outweigh sorry ones by at least 0.1 of the wave.Choice question so an illegal move is impossible, returned probabilities shade the pieces on the board, and a live calibration panel scores each claimed confidence against a one-ply material check.Boolean per legal cell in a single request, with no heuristic fallback and a user-set confidence threshold flagging unsure turns; the game is new, so there is no established play to copy, and its rules are not self-evident — they render from editable templates with auto-filled placeholders, so a designer can rewrite a rule and re-ask.Choice, so a button the game does not offer is impossible rather than unlikely, and 900 logged decisions at ~400 ms each on an M5 Pro never once answered off the menu.Choice decisions for StarCraft II macro control and micromanagement on 35 SMAC-Hard maps, validates selections against available actions, and follows optional GPT-6 plans to separate frequent action selection from longer-term strategy.Choice and four Score questions per fictional observer to animate 100 eyes from a shared post, revealing four amplified voices before equal-count analytics expose the full distribution of reactions.Choice over the 20 classic answers for the user's question and the page shows Jev's probability for every answer, displaying the highest-probability one because the rounded probabilities occasionally disagree with the reported choice.Choice per step, optionally guided by an LLM-written objective and standing rules; 3-seed ablations compare Jev over macro and raw actions against random, an LLM choosing every step, and Jev plus the planner.Choice questions (action, category, speed limit) about the caption and its lane or sidewalk, with no rule table overriding the answer; live at drive.mrza.ch, about 330 ms median and $0.00004 per decision.Score questions (move, turn) and one Noul (jump), with code acting on a score only past a 0.33 dead zone, jumping above 0.45, holding still any axis the current step does not use, and sending no Jev request while no step is active.Choice among the remaining route IDs while the server rejects any answer outside the supplied set, with self-hosted Laya using the same decision contract for comparison.Choice over named (emoji, square) pairs next to it and a Noul on whether the stroke is an unfinished shape, finishing the loop or line above 0.7 and otherwise sampling its pick from the returned probabilities (source).Source file: categories/robotics-physical.md
Choice decisions over structured state to select intent and Cartesian motion/gripper commands for a Franka Panda in MuJoCo, rejecting malformed responses and checking task success independently through physics.Source file: categories/finance-trading.md
Choice between buy and sell every block, publishing 81 ms decision latency and $0.000004 of Jev cost per call from a live dry run.Source file: categories/compliance-legal.md
Source file: categories/content-moderation.md
Boolean questions, benchmarked against TF-IDF baselines.Boolean "must this message be blocked?" plus a category Choice in one request, aborting the turn at 0.7 and failing open behind a deadline and circuit breaker; in production it blocked 9/9 hostile and 0/49 real messages at ~0.4 s median, about 4× cheaper than an LLM moderator.Choice per message in batches of 20, and shows a second column of only the messages matching a chosen intent (helpful, questions, funny, feedback); about 504 input tokens per message, roughly $0.15 per hour on a 2-message-per-second chat and $0.76 per hour at 50 per second.Noul for literal profanity in text or usernames and a second Noul for phonetic or look-alike disguise (a55h0le, mike_hunt); the threshold, max() policy, JSON response, and OpenAPI schema live in Worker code and the endpoint is callable from other Workers via service bindings.Choice (slop / not_slop) per X and LinkedIn post as it scrolls into view, blurring and stamping anything at or above a user-set threshold (default 0.7) behind a "Show the post" override, with a three-request concurrency cap and one cached verdict per post so scrolling never blocks.noul, choice, and score as first-class question types alongside a deterministic mock mode and three tests.Source file: categories/related-practices-discussions.md
Truncated — view the full README on GitHub.
85 followers · starred Sep 2026
734 followers · starred Sep 2026
1,258 followers · starred Sep 2026
4,841 followers · starred Sep 2026
Python
83.2%
Shell
16.8%
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
Python
2,061
678 commits
updated Oct 1, 2026
A curated awesome list of public projects and practices built on Jev, TypeSafe AI's System One model for typed decisions.
This README is the homepage aggregate of the current category files, so the latest accepted entries are visible here without drilling into subpages.
Jev is not a chat model. It takes unstructured state plus a typed question and returns a typed decision — a choice, a score, or a boolean, each with a confidence. That makes it a drop-in decision layer for software: classification, routing, rubric scoring, verification, and agent guardrails. This list tracks who is actually building with it, and which patterns transfer across industries.
The repository treats all categories equally — each entry lives in exactly one category, chosen by its direct Jev application domain. A dedicated Related Practices / Discussions category captures credible public practice signals — X threads, Reddit discussions, and interviews — that describe real Jev usage even when no strong standalone case page exists yet.
[!WARNING] A listing is not an endorsement. This project applies inclusion rules only — public, citable, genuinely uses Jev for a typed decision, one-sentence summary. It does not review code quality, security, maturity, or whether a project runs at all.
Treat same-day bulk submissions with particular care. Several repositories published together by one author, sharing a scaffold and a thin commit history, can satisfy every inclusion rule and still be unproven. Volume is not evidence of quality. See Curation is not endorsement for a checklist to run before adopting anything here.
Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. This list answers two practical questions quickly:
This is not a comprehensive database. It is a high-signal, fast-scanning field guide.
An entry should meet all of the following:
Jev/jev, cites TypeSafe AI's System One models, or shows a typed-decision loop (typed question → typed answer with confidence → accept/reject/escalate).We do not include:
Inclusion means one thing: the entry satisfies the inclusion rules above. It is not a quality review, a security audit, or a recommendation. We do not verify that a project compiles, that its tests pass, that its published numbers reproduce, or that its license permits your use.
This matters most for projects that arrive in bulk. When one author releases several repositories on the same day, they commonly share a single scaffold — the same AGENTS.md, CLAUDE.md, STATE.md, and CHANGELOG.md — land in one or two commits each, and may ship considerably more prose than code. Such projects can be entirely legitimate; they are simply unproven. Treat them as leads, not as validated tools.
Before adopting an entry, check it yourself:
| Check | Why it matters |
|---|---|
| Does the code actually call the Jev API? | An entry can read well on a README alone. Look for a real request carrying typed questions, and a parsed answer coming back. |
| Is there a runnable check? | A test, an example with expected output, or a public demo. No check means no evidence that it works. |
| Do the numbers have a source? | Any accuracy, latency, cost, or volume figure should be traceable to the linked page. We strip claims we cannot verify, but the project page itself may still carry them. |
| How much of the repository is code? | Some projects are mostly prompt documents. That can be legitimate — just know which one you are getting. |
| Is there a license? | A few entries have none, which limits reuse and redistribution. |
Found something wrong? Open an issue or a pull request — removal is as valid a contribution as addition. Rules for AI-assisted work, project depth, and submission rate live in CONTRIBUTING.md.
Each entry lives in exactly one category. When a project could fit multiple categories, we choose the one closest to its direct application domain.
Optional tags on an entry name the coding agent it targets and the kind of integration it is. Most entries carry none — they are added only when the source itself supports the classification.
Source file: categories/classification-routing.md
NOTRA_JEV_CLASSIFIERS flag routes brand-visibility classifiers off an LLM and onto Jev Boolean decisions at a 0.5 threshold, targeting 300 ms p50.Choice (invoice or general) and routes each inbound email to the matching handler./v1/systemone endpoint to new-api so typed decisions sit behind the same gateway as chat models.Choice and Noul questions, sending low-confidence answers to human review.Choice picks plain code, Jev or a reasoning LLM behind a Noul gate for non-tasks, code vetoes Jev when the idea needs images, and low confidence returns "not sure"; closed source, free page and API.Choice over the installed skill roster plus Boolean-style gates on whether any skill is needed, suggests a skill only when the gate and the per-candidate fit both clear 0.30, and defaults to a shadow mode that logs the decision without injecting it.Choice over ten kinds of post plus three Noul questions (paid ad, clickbait, emotional pressure) about each, counting an ad from 0.7, or from 0.4 when the kind is also ad, and clickbait and pressure from 0.5, then draws the monthly mix on a shareable card that links the highest-scoring posts for a manual check.Choice questions to classify CSV columns into a 13-code type vocabulary and each dataset into one of six scenes, then executes every write locally; measured Jev at 6.6–12.7× an LLM's token cost on this task because the output is already one character while per-question criteria repeat.jevonian/auto from session state, quota health, candidate capabilities, and cache-switch penalties, after deterministic code has filtered candidates and while pinned models, explicit jevonian/<route> requests, and routing.mode: "off" skip Jev entirely; minConfidence marks a low-confidence route in the ledger rather than accepting it, and the ledger records the serving model and why.Choice question per tab against user-editable group criteria — with unclassified tabs falling to a fixed fallback bucket, manual groups untouched, and the pre-grouping tab order restored on ungroup; an optional LLM engine with self-invented group names is included for comparison runs.Noul judgments against per-platform, user-defined topic and expression labels to annotate Weibo, Threads and X posts directly in a Chrome extension.Noul per remaining (source, destination) pair plus a guard Noul per incoming column, mapping 10 of 10 columns of a 23-column export at 253 questions in one call, 915 ms, $0.0012, unmapped fields left visible above a 0.75 threshold rather than guessed.Noul actionable, Score severity, Choice owning team, Choice duplicate-of), pages on P(SEV1)+P(SEV2) ≥ 0.80 and drops only below 0.20 when actionability agrees, sends the band between to a human who has 15 minutes to ack before it pages anyway, links duplicates into a cycle-broken incident graph so two alerts can never silence each other, and falls back to configured severity on any error or timeout; a published 300-alert run measured p50 418 ms, p95 1477 ms and $0.04 per 1,000 alerts, with the slowest call landing 151 ms short of the 2 s timeout.Choice and Noul questions about issue category and reproduction steps in opt-in live mode, then applies confidence and reproduction gates to propose a queue or review fallback without assigning the issue, with synthetic offline fixtures and policy tests.Choice, Score and Boolean questions to sort the peer's message into intent (10 options), an emotion distribution (9 options), urgency (0-3 Score) and a suggested reply posture (11 options), then shows exactly three toasts and takes no other action - no generated reply, no input injection, no screenshot or OCR, with a local alias-to-relation table passed as state so the same sentence is judged differently for a partner than for a colleague.Choice over member descriptions, fills blank task fields with Choice and Score questions, and routes automation workflows on a Choice/Score/Noul condition node, applying answers only at 0.6 confidence or above and otherwise leaving the task unassigned or taking the Else branch.Choice per user rule plus a Noul on whether the page is a payment, login or banking screen, stepping in with a pop-up only when a rule's probability clears its threshold and never on sensitive pages; on 119 trial pages with Kev, the short-video, feed, livestream and video rules had precision 1.00.Choice over the readings of a Chinese polyphonic character while the model stays fixed and only the harness around it is optimised, ending at half the errors for a seventh of the cost.Choice over /effort levels (low / medium / high / max / unclear) plus a Noul on whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set.Noul on the target plus Score rubrics about each row's text, turns every option's probability into a column next to the row's numeric fields, and lets a tabular foundation model such as TabPFN learn from the labeled rows in context, reaching 0.745 AUC at 256 labels on Kickstarter funding against 0.682 for Jev alone with calibration.Choice questions for kind, project and next action plus a Noul for duplicates, routing to a project only at 0.45 or above and marking a duplicate only at 0.70 with a specific matching item, while the text itself is stored verbatim in append-only local files.Source file: categories/adaptive-realtime-ui.md
Noul question, batched 20 at a time, answers whether it reveals a concrete plot event of the video being watched or of another title the user protects, with the extension owning the 0.85 / 0.7 / 0.5 threshold and keeping the comment covered when the check fails.Noul per menu item against the user's plain-language request, and presses the top match when it clears a probability threshold, falling back to a ranked list otherwise and never auto-running destructive items.Source file: categories/verification-guardrails.md
risk-check provider where Jev scores agent counterparties as typed Noul/Choice/Score questions into a code-controlled 0-100 score, issuing an ES256-signed attestation per verdict; 540-call scale run (99.76% at threshold 65-75, 0 false positives) and a 5-iteration 1,500-case adversarial red-team loop (100% adversarial accuracy) with ~$0.00005/decision at p50 ~400ms.Choice over the cross-turn coding-agent trajectory decides continue, pause, or escalate, with low-confidence verdicts routed to a human while deterministic code keeps dangerous-command blocking, budget caps, and an Ed25519-signed receipt chain.Noul checks about source and build files, escalates suspicious chunks for a second pass, and returns implicated files and lines before execution.typesafe_permission_reviewer builtin so the agent's permission decisions run through Jev rather than an LLM call.Boolean questions per paragraph (stacked hedges, restating closers, not-X-but-Y turns, naked cost figures) at a 0.7 threshold; CLI, pre-commit hook, GitHub Action and Claude Code skill; measured 182 ms median and 1 of 54 clean paragraphs flagged against 37 for Haiku 4.5.jev-pref.json rules that Jev checks against each diff hunk, staged file set, or pull request, returning fix_now or advisory findings to the coding agent and a nonzero exit code on blocking ones.Noul risk checks plus a risk Score, maps the answers in code to allow, ask or block by the user's risk setting (protected paths always ask), and also uses Jev to pick the model tier per prompt and to send back "done" claims that ran no verification, at about 400 ms per decision.Noul checks per tool call (beyond scope, destructive) and admits each judgment as Evidence that can degrade Trust, Delegation, and Authorization until a human override repairs the relation, so later calls inherit the history; includes a stateless-vs-stateful comparison with a scenario adversarial to persistence.Noul review question for each translated subtitle line to flag omissions, changed meaning, names, numbers, negation, or other defects for human review before export.Choice with an explicit Unresolved option to distinguish a mentioned arrival time from an agreed meeting time, showing the returned probabilities and prompting for missing agreement before treating a time as settled.Noul+Choice call per fact while deterministic code keeps the confidence gate; 90-sample calibration reports 0.922 gated accuracy and 0/15 injection flips.typesafe-ai/jev in malicious-link-check.ts before a short link is created, so the URL is gated by a typed verdict rather than a blocklist.Noul, Choice and Score questions about one function, file outline, candidate copy pair or test at a time, turns answers at 0.80 into review or consider findings with file and line, fails the build on review, and keeps undecided files as uncertain instead of clearing them.Choice risk class plus Noul irreversibility, task-match, and injection checks gate every non-read-only tool call (deny confident high-risk, ask ambiguous, delegate the rest), Noul scope and reversibility questions auto-approve clearly-granted calls behind argument-evidence gating, and a per-message Score re-ranks what a referenced session keeps instead of oldest-first dropping — fail-closed to stock behavior, shadow mode with a /jev-stats command, 64 tests.jev-guardrails mod that screens prompts and turns against jailbreaks, harm, and policy breaches via TypeSafe System One.Score lenses and one Noul — and blocks a diff, commit message or doc set that falls below the resulting quality score, from the terminal, a pre-commit hook or a GitHub Action.Noul judgments and local thresholds to flag team-rule violations as Claude Code and Codex edit, helping agents fix them before code review with configurable rule packs and repository-specific rules.Noul on whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable; 6 of 6 secrets and 0 of 6 benign strings were blocked in its published calibration.Noul for whether it has a bug, a Choice for which kind and which line, and a Score for severity, plus language-filtered CWE Noul checks and custom rules written as sentences at repository, file or method level, failing CI on any answer over its floor.Noul for each piece of code a rule applies to and reporting it above the rule's threshold; its two shipped rules were right on 12 of 12 sampled findings in three open-source projects.Score plus Noul checks for answer evidence and prompt injection, then withholds any draft whose claims fail a batched per-claim Noul or cite numbers absent from the sources; on its replayed 62-question golden set it answered 0 of 22 unanswerable questions, against 4 of 22 for naive top-5 RAG.Source file: categories/scoring-ranking.md
Noul to score already-loaded X posts against a reader's goal and profile, then pairwise Choice judgments to rank eligible matches and surface up to three for review.Score axes inside a single systemOne request and turns them into a 0-100 Slop Score in ordinary TypeScript.Noul properties per source file so the agent knows what to fix first.Noul checks to rank candidates, apply evidence thresholds, or extract source text while keeping those decisions independent.Noul judgments to find code, docs, logs, and text satisfying natural-language conditions, with a configurable probability threshold and ranked file results linked to source lines.Noul per line on whether it answers the query (16 lines per request, sharing the page text as state), highlighting lines at or above 0.55 page by page and ranking them by probability.Choice to locate source lines and Noul to check evidence presence across document regions in parallel, then returns original passages to a GPT research agent through Pi-Serini with 20/40/60-document batch limits.Noul relevance question per chunk across three parallel requests, and returns the accepted ones as excerpts through MCP; the cutoff and excerpt selection live in code.Score per item per weighted dimension of a YAML spec in a single request and sums weight times score in code to order the items, with noul, choice and score commands that turn a threshold into exit code 10 for shell scripts and CI.Score over each experiment result to decide whether it clears the promotion bar, and a Choice over candidate plays to decide what to run next in SEO, content, and ads.jev-latest against api.typesafe.ai and treats the returned probability as relevance, because TypeSafe exposes no native rerank endpoint.Noul per page on whether the visitor would be glad to land there, a Choice for the single best answer, and a Noul on whether any page answers at all) and re-orders or drops hits in code, with the repo's own benchmark over the 109-page TypeSafe docs reporting Hit@1 of 83% against 41% for its keyword pass alone.jevr) that turns a natural-language query into grep-style path:start-end targets — a stateless local BM25 pass proposes candidates, Jev Noul membership questions score their 100/20-line windows (kept at 0.90 for code, 0.60 for docs), and one listwise Choice per lane orders the keepers — ships as a Claude Code skill and plugin, and placed 2nd of 90 models on the HAKARI-Bench NanoRTEB reranking leaderboard.priorityGap that a local uncertainty policy turns into the order a human should read the hunks in, while OpenAI explains the ones that surface.Source file: categories/agent-decisions.md
Choice (allow / ask / block) with Noul irreversibility and exfiltration checks before every tool call, sends ask verdicts to a human and fails closed on errors, and adds a Choice model router with a confidence fallback and a Noul "am I done?" gate, each measured against labeled fixtures.Choice at each step to select a concrete socai CLI operation and observed post or profile target on Instagram, TikTok, or LinkedIn, rejecting malformed or low-confidence decisions before execution.Choice (wake, not yet, unrelated) against the agent's own sleep note, and code skips the wakeup only when wake is below 0.2 while always waking on user messages, bare timers, a skip limit, errors, and timeouts; one run passed 21 of 21 hand-written scenarios, which the README calls a smoke test rather than a benchmark.chrome_act_toward_goal) with an 85%+ pruned DOM tree (Shadow DOM & iframe pierced), dispatching native CDP events (isTrusted: true) on active logged-in sessions without focus theft.Choice picks the operation and DOM element each step and two Noul checks (goal reached, stuck) veto a premature DONE or BLOCKED, with a small text model used only when text must be typed.Choice per step (operation and target control) over the front window's accessibility table, executes only validated high-confidence answers, and escalates to a full LLM agent on low confidence, no-effect actions, or unknown field values; author-reported 275–690 ms per decision.Choice picks the next tool from candidates rebuilt every step, a Score grades the call's risk, and a Noul decides whether it needs authorisation, while plain code acts on the answers so a high risk score forces human authorisation that no probability can override (7.7% of wall clock with the offline judge, 79% over the hosted API).Choice decisions select tools and targets, uncertain decisions escalate to an LLM, and a shared guarded kernel supports isolated Docker workspaces and paired LLM-only comparisons.Choice over grid-tile candidates at each level, and local probability and margin gates decide whether to descend or refuse, emitting only a raster point and bounding box and never clicking.Choice over candidate replies, and fills the draft box while sending stays manual; a Windows port does the same from offline OCR of the WeChat window.jev-decide script and ships a P0 case set taken from accessibility-tree snapshots of a calculator, a calendar, and the NetEase home screen.jev_decide tool lets the agent ask a Noul, Choice, or Score question about any state — urgency triage, intent routing, guardrail checks — and gate on the returned probability or confidence in code instead of trusting the chat model's guess.Choice picks the operation and its target element each step from candidates read off the DOM rather than the layout, so the same agent runs on Chrome and on engines that never draw a page (Moli, Lightpanda, Kitesurf), passing at least 90% of runs on each of fourteen browsers tested.Choice, a Score and a Noul on each prompt and after every tool batch to pick the next API call's reasoning effort (low/medium/high), sent as a per-message statement so the prompt cache never breaks.Noul on whether a shell command is strictly read-only, auto-approving at 0.95 and otherwise falling back to the normal permission prompt without ever denying, while a local hard-no list and injection filter keep risky commands from reaching Jev; 0 of 8 state-changing commands were approved in its published calibration.Choice questions for the next operation and its target element behind the same /v1/systemone API, completing 38.5% of 125 hand-picked real-website tasks graded by deterministic verifiers versus 16.7% for Jev 1.13 in the same agent.Source file: categories/data-labeling-curation.md
Choice, Score, or Boolean decisions, sends ambiguous and audit samples to a human, and uses accepted human labels to optimize the saved definition with GEPA.tail -f output, by asking Jev one Noul per line against a plain-English question and printing lines at or above a probability threshold, with a hand-labelled benchmark against Claude in the repository.Choice per span (a schema class or none) plus a Noul per sentence-level class, keeping answers above a per-class threshold and flagging close calls for review, with a published benchmark measuring 10–26× lower cost than LangExtract on Gemini 3.5 Flash but lower F1 (84.2 vs 88.5 on its bilingual jx-bench).Choice over every candidate window per entity type, verifies each nominee with a second Choice (the type, none, mixed or partial) and settles its boundary with a third, averaging 73.7 strict F1 across 12 Chinese and English NER benchmarks against 72.1 for direct extraction with Qwen3.8-27B.Source file: categories/evaluation-benchmarking.md
Choice questions about first-visit understanding, returning inspectable findings for the first change to make.state costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing instructions in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.Noul rather than Choice moves the result more than the gap between the two models, at a thirty-fourth of the cost, written up at jjd-lab.github.io.Choice over the six faces of a hidden fair die 400 times; Jev selects face 1 on all 400 trials with 82.9% mean reported probability and 19.0% accuracy, then tests whether stated probabilities survive in synthetic forecast documents, where a 30% shortage risk comes back as 5.3% via Choice and 26.7% via Noul; raw responses and analysis code on GitHub.Choice (same / fact_differs / action_differs / specificity_differs) decides which of an agent's approved answers changed meaning rather than wording, and on 109 before/after pairs whose ground truth is derived from what each config rule does to the answer, Jev catches all 19 real changes with 13 false alarms against 33 for a markers-then-embeddings-then-LLM stack and 19 for the LLM judge alone.jev-1.13-20260917 through OpenRouter's TypeSafe-compatible /systemone endpoint, reporting about 261 fixed input tokens per request, zero spread in the implied per-request cost across question counts, batched-vs-single answer differences comparable to repeat-request noise, and median input-token savings of 76–86% at eight questions.Choice/Score/Noul or any OpenAI-compatible backend (with a free rules fallback), gates low confidence at 0.7 (caught 3/3 misjudgments at 9% escalation, n=130), and publishes Chinese-scenario cost-accuracy numbers — 97.7% @ ¥0.105/1k decisions and 60.0% → 68.3% on a frozen 120-item human-labeled spam set at τ=0.10.Choice on 108 verified four-option questions about the Convex backend platform (no docs or tools in the prompt, each asked 3 times with shuffled options, random guessing 25%) alongside 14 LLMs, where jev-1.13 scores 84.6% at a 199 ms median and $0.0088 per full run against 98.0% at 2.12 s and $1.59 for the top model, with every answer, probability and raw request/response in a public explorer and the runner in get-convex/convex-evals.Noul on whether an answer misrepresents its PubMed abstract; Jev scored 92.9% against 92.4–95.1% for four fast LLMs at a 204 ms median and USD 0.03 per 1,000 checks, and letting Jev settle the 37% of items where it was at least 90% sure kept each LLM's accuracy with 37% fewer LLM calls.Choice(2) up to Choice(16).Source file: categories/calibration-research.md
choice, score, and noul questions with RLCD-trained calibrated probabilities in a single ~35 ms forward pass, published on PyPI and Hugging Face.POST /v1/systemone server answering typed Choice/Score/Noul questions with confidence from open weights, verified as an official-SDK drop-in with temperature-fit calibration (set3 n=1316, 0.83 overall)./v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.POST /v1/systemone server that reads typed Choice/Score/Noul answers from the logits of any MLX checkpoint or OpenAI-compatible endpoint with no training, works as a drop-in for the official SDK, and replays Jev's published answers on 256 public judgments (231 vs Jev's 238, McNemar p = 0.21).Choice (2-255 candidates), Boolean, and Score questions with dev-calibrated distributions in one forward pass and zero output tokens, trained from scratch on CPU with committed datasets, predictions, and ECE results (maze 0.016).Choice/Score/Noul interface on commodity zero-shot NLI models and makes the confidence honest with temperature scaling and conformal abstention, shipping a reproducible calibration eval (ECE 0.170 to 0.071, cross-validated) that runs offline with no API key.Choice, Noul, and Score questions from structured state and documents without a hosted call.Choice, Score, and Noul in a single forward pass and returns calibrated confidence meant to be thresholded, so cases it is unsure about escalate to a human instead of being guessed; MLX-first on Apple Silicon, with a System One endpoint and weights on Hugging Face and ModelScope.Choice over 13 options (0–9, ., -, END) per answer character given the expression and the digits so far, appending whatever it picks and showing each step's full probability distribution and confidence, with a Rerun that exposes call-to-call jitter on identical inputs, live on Cloudflare Workers.Choice, Score, and Noul questions on the same POST /v1/systemone wire format, calibrated with temperature scaling plus a split conformal abstain set with a coverage guarantee (ECE 0.01 to 0.03 on the public suites), runs on CPU or in the browser via ONNX, fits on your own labels in seconds, and its README says it loses to Laya on typed decisions (0.71 vs 0.77).Choice questions one at a time, with confidence driving lookahead when the top pick falls below 0.65, beam search across sentence directions, and a self-critique loop that rewrites sentences scoring below threshold; live at talktojev.com, paper at doi.org/10.5281/zenodo.22940945.Noul, Choice and Score on the same POST /v1/systemone wire format in one forward pass, scoring 0.791 top-1 on typed-decisions' 2,000 test decisions after training on its train split (Jev 1.13.0: 0.727), and 0.774 on 873 scienthoon support tickets it never trained on (Jev: 0.753).Choice and Noul questions behind a TypeSafe-compatible API, matching Jev across computer-use, gaming and tool-calling with up to 9x lower latency and reporting 87.6% on Terminal-Bench 2.1 as a fine-tuned verifier.Choice, boolean or rubric-score question, released with its data pipeline, training config and eval harness under Apache-2.0, and reporting 90.1% on its 324-example holdout against Jev's 93.2%.Choice, Noul and Score questions with a probability for every option from one forward pass, and run anywhere: a standalone server with Jev's /v1/systemone wire and a playground, a llama.cpp script for the GGUF builds, or AINode (open-source local AI platform).Noul yes/no questions on Jev's own /v1/systemone wire format, running CPU-only via llama.cpp at 54 ms short / 220 ms long on a laptop Core Ultra 7 255H against Jev's hosted 344/345 ms; Choice and Score return 422, and it scores 0.815 accuracy on 2,000 unseen policy yes/no questions against Jev's 0.927.Choice, Noul and Score questions via prefill-only inference on NeoHorse-1-4B, deployable with vLLM, SGLang or a native Python/CLI/HTTP runtime, scoring 77.70 across six text benchmark groups (highest among open-weight entries with complete results in its published comparison).choice, noul and score on the same /v1/systemone format at about 22 ms per decision, published with a panel that measures Jev itself at 0.828 accuracy and 0.053 ECE while stating it claims no statistical significance.choice, noul and score on /v1/systemone with per-checkpoint provenance and calibration plots released./score call per step and supports per-task fine-tuning, released under MIT with the terminal run recorded.Source file: categories/infra-sdks-integrations.md
typesafe-ai/jev) in its experimental evaluate path.evaluate command.Choice, Score and Noul questions on Huawei 910B hardware — 37–47 ms median for a four-question request, 33.8x–70.9x faster than the same request on a single container CPU thread, with an output-equivalent SDPA decision head that avoids torch_npu's CPU fallback on aten::_transformer_encoder_layer_fwd.Scores each retrieved passage and Choice/Noul selects the query engine, with nfcorpus nDCG@5 0.340→0.396 at about $0.0003/query.npx skills add typesafe-ai/skills) that teaches agents the Jev workflow.jev-browser), so Jev arrives as a first-class Cline capability.Choice, Score and Yes/No questions to the model about pasted text - ticket triage, moderation, review scoring - and returns a parsed answer with a confidence value in about 0.5 s per decision.jev() filters, jev_prob sorts, and jev_choice groups.Choice/Score/Noul question sets with 13 offline lint rules before any Jev call, then sends the canonical wire payload and prints parsed, confidence-bearing JSON answers to stdout using exit code 2 to reject a billed-but-useless request.Noul, Choice and Score answers into named decisions with enter/exit thresholds (hysteresis), nested decision trees settled in one call, a JSONL journal, replay of a threshold change over recorded answers with no inference, and Brier/reliability calibration, over TypeSafe direct, OpenRouter or Vercel AI Gateway.Choice, Score and Noul questions through a decision client separate from the text model, the Kernel picks a specialist subagent by Choice, and each retrieved RAG passage is withheld from the generating model when its prompt-injection Noul exceeds 0.70.if Hunch.likely?("fraudulent", given: order) reads like plain Ruby but branches on a typed Jev answer, with pick for Choice, rate for Score, and graded predicates from possibly? to almost_certainly?; a TypeScript port offers the same interface.decide as a first-class inference type alongside generate and stream.Choice, Boolean, and Score decisions on pinned open models across Torch, vLLM, MLX, llama.cpp, and WebGPU, with committed row-level benchmarks and checksums.Choice, Noul or Score answer becomes a typed branch under caller-supplied thresholds, anything below them takes an Uncertain case the compiler forces you to handle, and procedure routing skips the model call entirely when deterministic predicates leave one candidate.Noul, Score, and Choice decisions into deterministic threshold workflows that batch into a single systemOne call and return an ordered, explainable action set instead of side effects, with matched rules recording the actual value behind each action and a mock provider so policy tests run without an API key.Choice, Score and Noul decision, trains a head per decision site, and answers live with calibrated confidence, falling back to the upstream below its threshold and demoting itself on drift; measured 22 ms per decision and 90% of support-triage questions answered locally at a 0.99 agreement target.Choice, Score, and Boolean evaluations; recipes display answer probabilities separately from provider confidence and use deterministic local thresholds to pause uncertain routes for review, while every configuration exports as TypeScript.Noul, Choice, and Score service that shares one state prefill across isolated questions and, on its included 27-question Qwen3.5-9B/A100 fixture, reports 0.500 s versus 13.554 s for generated JSON.Noul, Choice and Score questions that reads Choice and Score answers back as the caller's own types, rejecting any label or level the question never offered, with retries, OpenRouter support, and eleven examples tested against an in-process fake of the API.POST /v1/systemone wire contract, each rule citing TypeSafe's docs, OpenAPI file or SDKs, and a suite that checks any Jev-compatible server against it (Choice probabilities keyed by option and summing to 1, Score equal to Σ i·p, 2–255 options, error shapes, answers that stay put when question ids or order change), finding 2 of the 8 most-starred open ports conformant, with a reference mock that breaks each rule on purpose and a proxy that fixes what can be fixed.Noul, Choice, and Score questions and returning named answers as pipeline-friendly properties.TypeSafeModel integration to map Pydantic schema fields into typed Jev System One questions with confidence scoring.#[JevNoul], #[JevChoice]), Workflow guards, and WebProfiler panels.pip install "jev-style[torch]" (or [mlx] on Apple silicon) serves an open 0.8B Qwen3.5 decision model behind a System One-compatible /v1/systemone API that answers Choice, Score, and Noul questions with calibrated probabilities in one pass over inputs up to 25,600 tokens (0.15–0.2 s per short request with MLX on an M1 Max, after the first call), and ships a Claude Code guard that turns four Noul checks and a risk Score into allow, ask, or deny through code-owned thresholds.grev 'is a vegan meal' menu.txt keeps the lines whose Jev Noul clears 0.5, sibling filters route by Choice and rank by Score, and the output is always your own input, never generated text.Retry-After delay when the service answers 429, refuses any request that would pass a per-key daily spend ceiling (USD 0.20 by default), and keeps an opt-in cache of answer probabilities under caller-chosen keys, so a repeated question over the same state is not paid for twice./v1/systemone endpoint, so an existing Jev client only has to point TYPESAFE_BASE_URL at the daemon, and reports its recommended model at 0.722 accuracy against Jev's 0.738 on typed decisions.WHERE jev(people, 'the name is European'), jev_prob and jev_choice each put one Jev judgment per row with calibrated probabilities./v1/systemone proxy and Pydantic AI client that records every Jev answer's probabilities and alerts, without labels, when jev-latest switches versions, a question's answers drift (chi-square-tested PSI), answers crowd a decision threshold, or estimated accuracy falls.Choice, Score and Noul next to parallel fan-out, a router and a guardrail in one codebase.Choice, Score and Noul questions onto POST /v1/systemone from a separate provider package, with a runnable ASP.NET ticket-triage sample that picks the owning team and escalates to a human at a normalized Score of 0.8, and 1,096 tests across net8.0 and net10.0 that run with no HTTP.Source file: categories/game-simulation.md
Choice judgments to answer lateral-thinking puzzle questions and assess proposed solutions, with application code requiring supported facts, a coherent explanation, and sufficient confidence before marking a puzzle solved.Choice between two names under land, water or air rules held in state, asking the champion against the next K challengers in a single request and discarding the speculative answers once the champion falls — 1,999 fights in about 16 s at roughly US$0.01.Choice over four directions with no heuristic fallback, gated by a user-set confidence threshold that pauses play for human review, with editable prompts and board rules, bring-your-own-key backends, and archive import/export.Score questions on who would care plus seven Noul moderation checks (0.5 keeps the text out of the public feed, 0.85 blocks it), batched Choice questions then return each persona's reaction in waves of 600, 1,500 and 3,000, and code sends the text to the next wave only while glad reactions outweigh sorry ones by at least 0.1 of the wave.Choice question so an illegal move is impossible, returned probabilities shade the pieces on the board, and a live calibration panel scores each claimed confidence against a one-ply material check.Boolean per legal cell in a single request, with no heuristic fallback and a user-set confidence threshold flagging unsure turns; the game is new, so there is no established play to copy, and its rules are not self-evident — they render from editable templates with auto-filled placeholders, so a designer can rewrite a rule and re-ask.Choice, so a button the game does not offer is impossible rather than unlikely, and 900 logged decisions at ~400 ms each on an M5 Pro never once answered off the menu.Choice decisions for StarCraft II macro control and micromanagement on 35 SMAC-Hard maps, validates selections against available actions, and follows optional GPT-6 plans to separate frequent action selection from longer-term strategy.Choice and four Score questions per fictional observer to animate 100 eyes from a shared post, revealing four amplified voices before equal-count analytics expose the full distribution of reactions.Choice over the 20 classic answers for the user's question and the page shows Jev's probability for every answer, displaying the highest-probability one because the rounded probabilities occasionally disagree with the reported choice.Choice per step, optionally guided by an LLM-written objective and standing rules; 3-seed ablations compare Jev over macro and raw actions against random, an LLM choosing every step, and Jev plus the planner.Choice questions (action, category, speed limit) about the caption and its lane or sidewalk, with no rule table overriding the answer; live at drive.mrza.ch, about 330 ms median and $0.00004 per decision.Score questions (move, turn) and one Noul (jump), with code acting on a score only past a 0.33 dead zone, jumping above 0.45, holding still any axis the current step does not use, and sending no Jev request while no step is active.Choice among the remaining route IDs while the server rejects any answer outside the supplied set, with self-hosted Laya using the same decision contract for comparison.Choice over named (emoji, square) pairs next to it and a Noul on whether the stroke is an unfinished shape, finishing the loop or line above 0.7 and otherwise sampling its pick from the returned probabilities (source).Source file: categories/robotics-physical.md
Choice decisions over structured state to select intent and Cartesian motion/gripper commands for a Franka Panda in MuJoCo, rejecting malformed responses and checking task success independently through physics.Source file: categories/finance-trading.md
Choice between buy and sell every block, publishing 81 ms decision latency and $0.000004 of Jev cost per call from a live dry run.Source file: categories/compliance-legal.md
Source file: categories/content-moderation.md
Boolean questions, benchmarked against TF-IDF baselines.Boolean "must this message be blocked?" plus a category Choice in one request, aborting the turn at 0.7 and failing open behind a deadline and circuit breaker; in production it blocked 9/9 hostile and 0/49 real messages at ~0.4 s median, about 4× cheaper than an LLM moderator.Choice per message in batches of 20, and shows a second column of only the messages matching a chosen intent (helpful, questions, funny, feedback); about 504 input tokens per message, roughly $0.15 per hour on a 2-message-per-second chat and $0.76 per hour at 50 per second.Noul for literal profanity in text or usernames and a second Noul for phonetic or look-alike disguise (a55h0le, mike_hunt); the threshold, max() policy, JSON response, and OpenAPI schema live in Worker code and the endpoint is callable from other Workers via service bindings.Choice (slop / not_slop) per X and LinkedIn post as it scrolls into view, blurring and stamping anything at or above a user-set threshold (default 0.7) behind a "Show the post" override, with a three-request concurrency cap and one cached verdict per post so scrolling never blocks.noul, choice, and score as first-class question types alongside a deterministic mock mode and three tests.Source file: categories/related-practices-discussions.md
Truncated — view the full README on GitHub.
85 followers · starred Sep 2026
734 followers · starred Sep 2026
1,258 followers · starred Sep 2026
4,841 followers · starred Sep 2026
Python
83.2%
Shell
16.8%