v-modal/awesome-jev-tools

A curated list of tools built for Jev — TypeSafe AI's System One model for typed decisions.

754

96 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

List of tools around jev in production setup (r/LocalLLM)

Hello, Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. So, we gathered some tools/frameworks around jev: [awesome jev](https://github.com/v-modal/awesome-jev-tools) This list below answers two practical questions quickly: \- Where is Jev…

2

Oct 6, 2026

README

awesome-jev-tools

A curated awesome list of public projects and practices built on Jev, TypeSafe AI's System One model for typed decisions.

This README is the homepage aggregate of the current category files, so the latest accepted entries are visible here without drilling into subpages.

A curated list of public projects and developer patterns built on Jev, TypeSafe AI's System One model for typed decisions.

What is Jev? Jev is not a chat model. It does not write text or hold conversations.

Instead, it takes unstructured state alongside a typed question and returns a typed decision—such as a choice, a score, or a boolean—accompanied by a confidence rating.By eliminating token-by token decoding, Jev acts as a fast, low-latency decision layer directly inside software.

Developers use it to handle classification, infrastructure routing, rubric scoring, verification gates, and autonomous agent guardrails.Goal of this ListMost discussions about Jev are scattered across launch threads, social media, and one-off prototypes.

This repository centralizes those pieces to answer two practical questions for developers:

  • Production Validation: Where is Jev actively making real decisions in live production workflows?
  • Transferable Patterns: Which decision architectures can be cleanly copied and applied across different industries?

Goal of this list

Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. This list answers two practical questions quickly:

  • Where is Jev already making real decisions in production workflows?
  • Which decision patterns transfer across industries?

Inclusion criteria

We do not include:

  • Generic classifiers, routers, or research agents that merely resemble the pattern without using Jev.
  • Pure theory or opinion without a concrete practice.
  • Launch-hype commentary with no working artifact or reproducible result.
  • Long write-ups inside the list itself.
  • Sources that are private, inaccessible, or too vague to classify.

Curation is not endorsement

Inclusion means one thing: the entry satisfies the inclusion rules above. It is not a quality review, a security audit, or a recommendation. We do not verify that a project compiles, that its tests pass, that its published numbers reproduce, or that its license permits your use.

This matters most for projects that arrive in bulk. When one author releases several repositories on the same day, they commonly share a single scaffold — the same AGENTS.md, CLAUDE.md, STATE.md, and CHANGELOG.md — land in one or two commits each, and may ship considerably more prose than code. Such projects can be entirely legitimate; they are simply unproven. Treat them as leads, not as validated tools.

Before adopting an entry, check it yourself:

CheckWhy it matters
Does the code actually call the Jev API?An entry can read well on a README alone. Look for a real request carrying typed questions, and a parsed answer coming back.
Is there a runnable check?A test, an example with expected output, or a public demo. No check means no evidence that it works.
Do the numbers have a source?Any accuracy, latency, cost, or volume figure should be traceable to the linked page. We strip claims we cannot verify, but the project page itself may still carry them.
How much of the repository is code?Some projects are mostly prompt documents. That can be legitimate — just know which one you are getting.
Is there a license?A few entries have none, which limits reuse and redistribution.

Found something wrong? Open an issue or a pull request — removal is as valid a contribution as addition. Rules for AI-assisted work, project depth, and submission rate live in CONTRIBUTING.md.

Current coverage

Each entry lives in exactly one category. When a project could fit multiple categories, we choose the one closest to its direct application domain. You can browse the list by category below.

Open categories still being seeded

  • Scientific Pipelines — 0 entries

Full list

Classification & Routing

Source file: categories/classification-routing.md

  • JEV Book Tags - Library cataloguing: a calibre plugin asks Jev Noul questions about book genres and subjects, applies configurable per-tag probability thresholds, and preserves existing tags while leaving uncertain results for review.

  • Notra - Marketing analytics: production GEO platform whose NOTRA_JEV_CLASSIFIERS flag routes brand-visibility classifiers off an LLM and onto Jev Boolean decisions at a 0.5 threshold, targeting 300 ms p50.

  • jev-router - Developer tooling: routes Claude Code tasks to the cheapest capable model by asking Jev to choose among candidates.

  • Codex Jev Router - Coding agents: asks Jev Choice and Noul questions to select Codex subagent model and reasoning tiers, with confidence gates and a Sol fallback.

  • jev-router (prismhq) - LLM infrastructure: open-source LiteLLM-based router where a Jev decision picks which model serves each request.

  • pi-jev-router - Coding agents: adds automatic per-request model routing to the Pi coding agent through Jev decisions on Vercel AI Gateway.

  • jcm-router - Coding agents: local proxy that picks the Claude model and reasoning effort per message with a Jev decision while leaving the cached main chat untouched.

  • Switchboard - Developer tooling: assesses a new Claude Code or Codex conversation with Jev, then local confidence policy selects and pins its model and reasoning effort through follow-ups, tool calls, and resume to avoid unnecessary prompt-cache disruption.

  • jev-agent-skill-router - Agent infrastructure: routes agent skill selection through typed, confidence-aware Jev decisions so weak matches are declined instead of guessed.

  • typesafe-jev CV screener - Recruiting: screens a folder of CVs with Jev typed judgments against an editable policy, re-scoring candidates for free when the policy changes.

  • Jev email intent workflow - Back-office automation: async LangGraph workflow gets a typed Jev Choice (invoice or general) and routes each inbound email to the matching handler.

  • unclutter - Browser tooling: WXT extension where Jev decides per page element whether it is clutter, removing it under reusable template rules.

  • typesafe-adblock - Browser tooling: Chrome extension that asks Jev whether each DOM element is an ad, turning ad blocking into a stream of per-element typed questions.

  • DiffJury - Code review: routes each pull request by risk with Jev before a human reviewer is assigned, doubling as a review coach.

  • HA-Jev - Smart home: Home Assistant integration that answers questions about the house as a probability, a choice, or a score.

  • secondlayer - Fault triage: self-hosted Stacks data service whose Slack gate and fault-triage paths both run on Jev decisions.

  • new-api-typesafe-plugin - LLM gateway: adds a native /v1/systemone endpoint to new-api so typed decisions sit behind the same gateway as chat models.

  • duet-agent - Agent harness: keeps a Jev-backed routing table for deciding which model should serve a request.

  • json-render - Generative UI: Vercel Labs' UI framework uses Jev in its compose path to pick which components and actions a rendered interface should contain.

  • omo-jevlike-router - Skill routing: shrinks the skill catalog in a system prompt with one forward pass over a frozen Qwen, routing each request Jev-style.

  • jev-cookbook - Developer education: 15 runnable Node recipes that route support tickets, file documents, categorize bank transactions and label Gmail with Jev Choice and Noul questions, sending low-confidence answers to human review.

  • flue-jev-demo - Agent routing: routes a Flue agent's work with Jev through Cloudflare AI Gateway.

  • sift - Content labelling: Chrome extension that labels every post in an X timeline - substance, humour, chit-chat, promo, junk, or AI-written - with Jev decisions.

  • jev-tree - Classification: A hierarchical selector that asks Jev to traverse branches when a catalog has too many options for one question.

  • litellm - Model Routing: LiteLLM can use Jev to classify requests for its complexity-based model router.

  • oh-my-pi - Model Routing: Oh My Pi includes an optional TypeSafe judgment provider for bounded decisions in coding-agent workflows.

  • jev-model-router - Model Routing: A community Claude Code Templates mod that uses Jev to suggest subagent models and reasoning levels.

  • openchamber - Model Routing: OpenChamber’s optional automatic model router uses Jev to classify a message before selecting a configured model and reasoning level.

  • firstmate - Model Routing: Firstmate can optionally use Jev to match task briefs to dispatch rules before local policy chooses an Agent configuration.

  • hermes-jev-skills - Model Routing: Jev-powered model routing, memory, compaction, skill selection, computer and browser use for Hermes agents (also Claude Code and Codex)

  • vexjoy-agent - Model Routing: An optional Jev routing path that matches VexJoy requests to specialist Agents, skills and workflows.

  • WrongStack - Model Routing: An optional Jev dispatch classifier for choosing among WrongStack specialist Agents.

  • jev-codex-router - Model Routing: Uses Jev to classify each Codex turn, then applies local rules to choose the model, reasoning effort, and speed mode.

  • JevRouter - Model Routing: Routes among models, subagents, skills, MCP tools and CLIs using a shared candidate set.

  • grok-bot-jev - Model Routing: Connect TypeSafe Jev to Grok Bot as a cheap decision layer - usage gates, skill template, examples

  • loki - Model Routing: Loki optionally adds Jev typed-judgment tools and routes a new session to a model within the selected gateway.

  • JevLoop - Model Routing: The agent loop where decisions don't cost a large language model call. Zero deps, runs offline, no API key needed.

  • muse-jev-playbook - Model Routing: Jev decision layer for Muse: a fast, cheap TypeSafe AI gate before expensive agent work — confidence policy, recipes, reference router, honest measurement.

  • sabi - Model Routing: Adaptive inference scheduling for AI agents — per-round model, effort and provider routing for coding harnesses: a Command Code mod or a local OpenAI-compatible proxy.

  • typesafe-skill-router - Model Routing: An opt-in Hermes Agent plugin that asks Jev to suggest one relevant skill before a model call.

  • dejevu - Model Routing: Jev? Déjà vu. Browser agents that run on instinct, no Jev needed. One look at the page, one call to any open model, one action. Faster than the Jev demo on Google Flights.

  • jev-router - Model Routing: Cost-aware LLM router that picks the cheapest model capable of handling a query, using TypeSafe's Jev for fast classification instead of an LLM call.

  • laya-jev-lab - Model Routing: Independent measurements of typed-decision models: Jev (TypeSafe API) vs Laya (open weights), and a local-first cascade that matches Jev's accuracy at 1.8x the speed

  • Jev-Auto-Router - Model Routing: Jev Auto Router (Jev Router): experimental per-call GPT model routing for Codex via TypeSafe Jev and a local Responses proxy, with independent task verification.

  • tool-prune - Model Routing: Calibrated tool selection and schema pruning for AI agents. Dual-engine: zero-dependency offline TurboQuant or TypeSafe System One (Jev). Prunes candidate MCP tools and schemas down to the relevant set before calling LLMs to eliminate hallucinations and save tokens.

  • jev-for-all - Model Routing: Jev for every agentic development workflow — the System One decision model wired into whatever harness an agent codes in: OpenCode today, Claude Code and Hermes adapters next.

  • Jev-Model-Router-Claude-Code - Model Routing: Begleitmaterial zum Video „Jev + Claude Code: 3 Use Cases".

  • jev-opus - Model Routing: Claude Opus 5.5 with the effort level re-decided every step by the TypeSafe Jev reflex — without breaking the prompt cache. CLI + Claude Code plugin.

  • jev-pilot - Model Routing: Let Jev steer Claude Code: the right reasoning effort, subagent model and skill for every prompt. A Claude Code plugin powered by TypeSafe's Jev (OpenRouter / TypeSafe).

  • jev-claw - Model Routing: Typed model routing for OpenClaw agents, powered by TypeSafe Jev

  • jev-model-router - Model Routing: Model router for Claude Code using Jev

  • jev-router - Model Routing: Pass-through model router for Claude Code and Codex CLI that picks a model tier per human turn with Jev, TypeSafe AI's decision model

  • jev-smart-router - Model Routing: JEV Smart Router — a Databricks App that uses TypeSafe JEV to pick which model answers each message, then runs inference on the chosen Databricks Foundation Model API endpoint.

  • opencode-jev-orchestrator - Model Routing: An OpenCode orchestrator that keeps a cheap sticky parent model and, when Jev flags a hard turn, escalates through a child subagent.

  • tiershift - Model Routing: Policy-driven model routing framework routing every LLM call to the cheapest capable tier in ~180 ms via TypeSafe Jev.

  • todo-jev - Model Routing: A task-routing experiment combining skill conditions and environment checks to suggest rules, skills or a large model.

  • chat2jev - Model Routing: Convert OpenAI-compatible Chat Completions requests into TypeSafe System One (Jev) **State / Questions**, compare generated text with structured judgments, and publish reusable question sets as proxy routes.

  • Janus - Model Routing: Framework for measuring when to employ Jev versus generative LLMs on proprietary datasets, routing queries based on measured benchmarks.

  • jev-codex-model-and-effort-router - Model Routing: Copy and paste this into your coding agent:

  • jev-codex-pilot - Model Routing: A Codex overlay incorporating JEV to make the best decisions regarding model selection and depth of reasoning. All while automating the process via an automated Kanban system.

  • jev-route - Model Routing: **Run it. Log it. Distill it. Own it.**

  • jev-routing-experiment - Model Routing: Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena

  • codex-jev-native-router - Model Routing: Experimental native Codex Desktop and CLI model routing with Jev and a configurable allowlist

  • hermes-jev - Model Routing: Jev decision sidekick for Hermes Agent — TypeSafe and Cloudflare, explicit tools and official skill

  • jev-agent-hooks - Model Routing: TypeSafe Jev hooks for Claude Code, Codex and pi: per-turn skill suggestion and subagent model routing

  • jev-claude-router - Model Routing: Model router for Claude Code using Jev

  • jev-model-router - Model Routing: Cost-optimized OpenRouter model router using TypeSafe's Jev, with a live full-catalog scorer instead of a hardcoded model list

  • jev-router-playground - Model Routing: A model-routing playground where Jev picks a candidate and users compare the resulting answers.

  • jevbus - Model Routing: A streaming event bus whose routing, subscription and consumption are decided by a probabilistic judge. The reference judge is TypeSafe AI's Jev (System One) model: send it a payload and a set of typed questions, get back calibrated probabilities instead of prose.

  • openclaw-jev-plugin - Model Routing: A silenced message never reaches the language model, so it costs one Jev call and no model tokens. Direct messages always get an answer unless you choose to gate them too.

  • stuntdouble - Model Routing: Drop-in /v1/systemone proxy that shadows Jev with local decision models (Kev, Laya) and reports whether you can swap

  • hermes-jev-router - Model Routing: Hermes Agent plugin: TypeSafe Jev model routing + trim-then-compress

  • jev-lab - Model Routing: Open lab: Jev (TypeSafe System One) routing in front of Claude Code - measured bugs, patch, and a hard fallback with alerts

  • jev-research - Model Routing: A Jev and Herdr integration guide with a prototype for routing tasks to different Agents.

  • JudgeJev - Model Routing: The public site replays **17 real, recorded Jev evaluations** of fictional support cases. It starts with a false shipping guarantee that scored **15.1%**, alongside a correct control that scored **95.5%**. You can inspect raw requests and responses, adjust decision rules, and see why a five-pair candidate fails a release gate.

  • omp-plugin-jev-router - Model Routing: Route Oh My Pi prompts between simple and advanced models with TypeSafe AI's Jev classifier.

  • jev-cc-codex-router - Model Routing: Per-turn model routing proxy for Codex: asks Jev which tier each task needs, rewrites the model, retries flaky upstream errors.

  • jev-decision-gateway - Model Routing: A gateway that asks Jev continue / tool / verify questions and invokes a generative LLM only when policy says generation is needed.

  • jev-demo - Model Routing: A customer-service routing demo that batches Jev questions before following the resulting route.

  • jev-gateway - Model Routing: Session-aware OpenAI-compatible model-routing gateway powered by JEV

  • Diffusion Jev - Visual classification: independent Jev-style DiffusionGemma/SGLang server that selects doodle and flower labels from image pixels with typed Choice questions and displays candidate scores in a drawing playground, with public evaluation artifacts and uncalibrated probabilities.

  • jev-logtriage - On-call operations: batches collapsed Loki logs into one Jev call of Noul, Score, and Choice questions, then maps answers in code to suppress, watch, review, notify, or page, with low confidence going to review and nothing executed.

  • DocJev - Document pipelines: LlamaIndex's open-source library that classifies a document against natural-language category rules or finds the boundaries between sub-documents, with swappable OCR backends (liteparse or LlamaParse) and a benchmark harness whose 40-document pilot classified 40/40 originals correctly at about 182 ms Jev decision p50.

  • jev-fit - Developer tooling: hosted fit checker that sends a pasted software idea and a fixed typed rubric to Jev in one call, where a Choice picks plain code, Jev or a reasoning LLM behind a Noul gate for non-tasks, code vetoes Jev when the idea needs images, and low confidence returns "not sure"; closed source, free page and API.

  • Jev-Mail - Email productivity: runs a 24/7 Gmail classifier on user-owned Google Apps Script where Jev scores urgency, importance, and category, routing uncertain or suspicious mail to Review without a local daemon.

  • AI-decision-maker - Data cleaning: asks Jev Choice questions to classify CSV columns into a 13-code type vocabulary and each dataset into one of six scenes, then executes every write locally; measured Jev at 6.6–12.7× an LLM's token cost on this task because the output is already one character while per-question criteria repeat.

  • hearth-jev-rental-search - Housing search: autonomous multi-source rental search where Jev decides which listings match the criteria.

  • pi-jev-skill-picker agent: Pi - Coding agents: ranks the Pi agent's installed skills against the current task with Jev before any of them run.

  • Feed Lens type: extension - Social media: uses Jev Noul judgments against per-platform, user-defined topic and expression labels to annotate Weibo, Threads and X posts directly in a Chrome extension.

  • jev-table-import-mapper - Data import: maps an uploaded CSV's columns onto a destination table with a strict deterministic name-equality pass, then one Jev Noul per remaining (source, destination) pair plus a guard Noul per incoming column, mapping 10 of 10 columns of a 23-column export at 253 questions in one call, 915 ms, $0.0012, unmapped fields left visible above a 0.75 threshold rather than guessed.

  • Jevidence - Developer education: Python sandbox asks Jev Choice and Noul questions about issue category and reproduction steps in opt-in live mode, then applies confidence and reproduction gates to propose a queue or review fallback without assigning the issue, with synthetic offline fixtures and policy tests.

  • langchain-skill-router - Agent infrastructure: per-turn skill routing for LangChain deepagents, where Jev ranks the SKILL.md catalog against the request and the recent conversation and verifies the top candidates, so only the picked skill's instructions reach the prompt; the judge is a protocol that a self-hosted model or static rules can implement instead.

  • jev-rental - Consumer rental: sorts every claim in a rental listing into verify-on-site / demand-evidence / high-risk-pitch buckets to build a pre-viewing checklist with code-templated questions; 50-sample calibration reports 0.910 gated accuracy and 0/10 injection flips.

  • jev-resume-disqualifier - Recruiting: knocks a resume out of a pipeline in under 25 ms by asking Jev the disqualifying question first, so only survivors reach a full evaluation.

  • Jev-IOT - Smart Utilities & Telecommunications: Ultra-low-cost, non-autoregressive AI telemetry classifier enabling sub-150ms anomaly triage and autonomic remediation across 10M+ smart meters for under $35/month.

  • AgentScope - Multi-agent platforms: multi-agent platform by Alibaba implementing native TypeSafe Jev classification models for binary, choice, and score routing across agent pipelines.

  • inbox-zero - Email productivity: open-source AI email assistant that uses TypeSafe Jev System One decision models to classify incoming email intent and triage action items.

  • SiYuan - Knowledge management: privacy-first personal knowledge management system featuring native Jev decision model integration for high-speed document classification, flashcard intent categorization, and automated tag routing.

  • Paca type: self-hosted - Project management: self-hosted open-source Jira alternative that auto-assigns tasks with a Jev Choice over member descriptions, fills blank task fields with Choice and Score questions, and routes automation workflows on a Choice/Score/Noul condition node, applying answers only at 0.6 confidence or above and otherwise leaving the task unassigned or taking the Else branch.

  • Qualm - Digital wellbeing: macOS menu bar app that reads the screen as text through the Accessibility API and asks Jev (or Kev, its local open-source counterpart) one Choice per user rule plus a Noul on whether the page is a payment, login or banking screen, stepping in with a pop-up only when a rule's probability clears its threshold and never on sensitive pages; on 119 trial pages with Kev, the short-video, feed, livestream and video rules had precision 1.00.

  • Auto-optimizing Jev: half the errors, 1/7 the cost - Text classification: asks Jev a Choice over the readings of a Chinese polyphonic character while the model stays fixed and only the harness around it is optimised, ending at half the errors for a seventh of the cost.

  • spending-effort-with-jev agent: Claude Code type: plugin - Coding agents: Claude Code plugin whose UserPromptSubmit hook asks Jev a Choice over /effort levels (low / medium / high / max / unclear) plus a Noul on whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set.

  • tab-jev type: library - Tabular prediction: asks Jev a Noul on the target plus Score rubrics about each row's text, turns every option's probability into a column next to the row's numeric fields, and lets a tabular foundation model such as TabPFN learn from the labeled rows in context, reaching 0.745 AUC at 256 labels on Kickstarter funding against 0.682 for Jev alone with calibration.

  • tinystruct-typesafe-sdk - SDK: TypeSafe Jev integration library for building type-safe classification and routing decisions with structured outputs.

  • sortwell agent: Claude Code type: plugin - Personal inbox: MCP server and Claude Code plugin that files each captured note, link or meeting line with one Jev request of Choice questions for kind, project and next action plus a Noul for duplicates, routing to a project only at 0.45 or above and marking a duplicate only at 0.70 with a specific matching item, while the text itself is stored verbatim in append-only local files.

  • IntentSQL - Natural-language SQL: turns a question about a SQLite database into a sequence of small Jev decisions instead of one generated query, released as an experiment alongside its decision lab.

  • TypeSafe Conversation - Home automation: a Home Assistant voice agent built on Jev.

  • Jevvie type: library - Web companions: a page offers its actions as WebMCP tools and one Jev Choice picks the action a visitor's request means, with a Choice per argument asked alongside, asking back when the top two options are close and gating unprompted tips with a Noul (source).

  • Gut Check - Smart home: Home Assistant integration whose eight install checks ask Jev a Score on each pending update's release notes and a Choice per item elsewhere, such as whether an unavailable entity is expected, worth fixing or safe to remove; answers below 0.5 confidence change nothing, and the rest that need action become Repairs cards the user must confirm.

Adaptive & Realtime UI

  • DWIM - Desktop productivity: a macOS command palette that reads the frontmost app's menu tree through the accessibility API, asks Jev one Noul per menu item against the user's plain-language request, and presses the top match when it clears a probability threshold, falling back to a ranked list otherwise and never auto-running destructive items.

  • shapeshift - Input: one text box that morphs into the right UI as you type, asking Jev which control the sentence calls for, and running offline.

  • Jevcast - Desktop productivity: native macOS launcher and window manager that uses Jev to match natural-language window and action commands to known application workflows with local response caching.

Verification & Guardrails

Source file: categories/verification-guardrails.md

  • GeekLink Jev Subtitle Translator - Media localization: reviews source–translation SRT pairs with one Jev Noul decision per cue and flags suspected omissions or meaning changes for human review.

  • is-malicious - Software supply-chain security: asks Jev Noul checks about source and build files, escalates suspicious chunks for a second pass, and returns implicated files and lines before execution.

  • jev-review - Software engineering: staged code-review workflow and local dashboard where Jev gates each review stage before a change advances.

  • pi-jev - Agent safety: adds a measured tool-call gate to the Pi coding agent so risky calls are checked by Jev before execution.

  • OpenWork - Engineering workflow: wires Jev into its eval testkit as a verification judge so agent-produced work is gated by typed verdicts rather than a text model.

  • jev-guard - Agent security: prompt-injection and dangerous-action guard for Claude Code, Codex, Pi, and ACP agents, with Jev deciding what to block.

  • Foreman - Software factory: sits above Codex workers and has Jev independently judge whether an implementation is complete, its tests sufficient, or a human is needed.

  • stanley-code - Coding agents: bounded Jev workflows that keep agent judgments typed instead of free-form.

  • opencompany - Agent workspace: runs its approval review through Jev so workspace actions are gated by a typed decision.

  • jev-git - Developer tooling: sub-second Git pre-commit & pre-push reflex gate that screens staged diffs for secrets and destructive commands using Jev.

  • pi-heed - Runtime constraints: checks every side-effecting tool call from the Pi agent against what the user actually asked for.

  • Hunch - Code review: plain-English rules that Jev checks code against, locally or on every pull request, with Jev picking one label per finding.

  • Abide - Agent supervision: reads every edit a coding agent makes and has Jev flag rule violations, with the project reporting that an independent reviewer confirmed 10 of the 39 flagged edits and 11 of the 15 flagged turns.

  • fx - Coding agent: ships a typesafe_permission_reviewer builtin so the agent's permission decisions run through Jev rather than an LLM call.

  • Sniff Test - Writing: prose linter that asks Jev ten Boolean questions per paragraph (stacked hedges, restating closers, not-X-but-Y turns, naked cost figures) at a 0.7 threshold; CLI, pre-commit hook, GitHub Action and Claude Code skill; measured 182 ms median and 1 of 54 clean paragraphs flagged against 37 for Haiku 4.5.

  • jev-pref - Code review: turns the preferences in a project's AGENTS.md into jev-pref.json rules that Jev checks against each diff hunk, staged file set, or pull request, returning fix_now or advisory findings to the coding agent and a nonzero exit code on blocking ones.

  • jev-axi - Agent safety: PreToolUse gate for Claude Code and Codex that has Jev score each shell command for destructiveness, exfiltration, remote code execution, and security weakening, deciding routine commands locally so nothing is sent for them, and scoring 44/44 on the 44 labeled tool calls in its repository.

  • pi-verdict - Agent safety: Pi permission gate where Jev answers one Choice (allow/ask/deny) per gray-zone tool call — deterministic rules settle clear cases first, deny blocks, ask escalates to a human confirm, and errors or timeouts deny; Jev is an optional backend, OpenRouter-only and experimental.

  • jev-commit - Developer tooling: pre-commit hook where one Jev call judges whether the commit message matches the staged diff, flags debug leftovers and unmentioned work, and blocks only on a detected credential.

  • Blink - Code review: CLI that coding agents run after every change, with Jev checking the diff near-instantly in place of an LLM reviewer.

  • hermes-jev-approvals - Agent approvals: proof of concept that puts Jev in front of Hermes Agent's command approvals, reporting 8.7x faster decisions and 4.4x fewer prompts to the user.

  • jev-engineering - Agent safety: gates coding-agent tool calls with deterministic rules first and one typed Jev call second, then publishes a rerunnable 300-call injection test showing what the gate catches and what walks past it.

  • Cribrix - Retrieval / RAG: filters retrieved chunks with a Jev Score plus Noul checks for answer evidence and prompt injection, then withholds any draft whose claims fail a batched per-claim Noul or cite numbers absent from the sources; on its replayed 62-question golden set it answered 0 of 22 unanswerable questions, against 4 of 22 for naive top-5 RAG.

  • agentgateway - Security & Guardrails: A Jev guardrail example in Agentgateway using a webhook to inspect model requests and responses.

  • Agent - Security & Guardrails: An optional Jev command-risk advisor inside a native macOS Agent, with a TypeSafeKit client.

  • interlinked-cli - Security & Guardrails: Interlinked adds optional Jev judgments and evidence checks to local coding-agent checks.

  • pi-warden - Security & Guardrails: Adds checks for project rules, out-of-scope actions, repeated failures, and completion claims to Pi Agents.

  • building-with-typesafe-jev - Security & Guardrails: Unofficial skill that teaches coding agents to build with TypeSafe AI's Jev: typed decisions, calibrated confidence, and prior art from 150+ community projects.

  • jevals - Security & Guardrails: Agent evals and guardrails as Jev decisions: one request per trace, a fraction of a cent, fast enough for the agent loop. Runs locally with Kev or Laya.

  • captaincore - Security & Guardrails: Jev commands in the WordPress toolkit CaptainCore answer structured questions and prioritize malware scanner findings for review.

  • jev-kit - Security & Guardrails: Everything you need to run TypeSafe's Jev with Claude Code: a tool-call guard, tier guard, file search, browser agent, review, belay, compaction and installers.

  • jev-edge - Security & Guardrails: Typed-judgment admission control at the traffic edge: three-layer prompt-injection and abuse filter for nginx/OpenResty, powered by TypeSafe Jev. Fail-open, cached, hot-reloadable.

  • pi-jev-auto-mode - Security & Guardrails: Adds rule checks to Pi commands and file operations, then uses Jev to assess cases that need further judgment.

  • jevvy - Security & Guardrails: Jev-powered plugins for coding agents

  • jev-safety-gateway - Security & Guardrails: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-security-scan - Security & Guardrails: Reviews Agent Skills and MCP configurations and source code with local static checks and TypeSafe Jev before installation or execution, reporting file and line evidence, risk categories, model probabilities, and coverage gaps.

  • jev-enforce - Security & Guardrails: 📏 Claude Code plugin that makes Claude follow your CLAUDE.md: every reply and edit checked by TypeSafe Jev ✅

  • jev-engineering - Security & Guardrails: Jev Engineering: Typed Decision Systems for Reliable Agent Workflows. Paper, diagrams, and companion examples by Av1dlive.

  • JevPR - Security & Guardrails: PR Risk review, automated by Jev

  • pi-jev-sentinel - Security & Guardrails: Open-source Pi coding-agent extension that uses Jev to check tool calls before they run, scan files for prompt injection, flag risky replies, and keep secrets out of what it sends.

  • dsh-jev - Security & Guardrails: Jev (System One decision model) plugin suite for DeepSeek Harness (dsh)

  • jev-auto-approve - Security & Guardrails: Jev is a decision model: it answers a typed question with a calibrated probability rather than prose. This action asks it one yes/no question per thing worth being sure about — answered in parallel in a single call — and approves only when every one of them clears your threshold:

  • jev-block-android-ad - Security & Guardrails: An Android notification and SMS filter that applies local OTP rules before asking Jev whether a message is advertising noise.

  • jev-guard - Security & Guardrails: Probability-scored guardrails for Claude Code: deny rule-breaking edits and unasked-for deploys, route your docs into each prompt, and check the final answer against the turn's own evidence.

  • jev-phishing-bench - Security & Guardrails: The signal result above was challenged on three points: no non-AI baseline, selection and evaluation on the same emails, and no equivalent decomposition for the LLM. Three controls were added (`bench/heuristics.py`, `bench/protocol.py`, `run_llm_signals.py`); nothing above was changed. Full tables in `results/report.md`, chart in `results/controls.png`.

  • jev-tool-permissions - Security & Guardrails: Adds tool-call approval and tool-list pruning to the Vercel AI SDK.

  • pi-jev-guard - Security & Guardrails: Check Pi code edits against repository Markdown rules with TypeSafe Jev

  • agi-jev-containment - Security & Guardrails: **Open-source AI agent monitoring, malicious-agent detection, and escalate-only containment** for sandboxed LLM agents. Local HackSpain 2026 stack (AngryRobot dashboard): FastAPI, React/Vite, Neo4j. Classifies a *chain of actions*, not a single tool call. A model never pulls the plug.

  • jev-baselines-eval - Security & Guardrails: Every latency number here is **wall-clock duration of one API call** measured in the client: a timestamp before the request, another after the full response body is read ([`code/run_b1.py:32`](code/run_b1.py), [`code/common.py:50`](code/common.py)). Non-streaming on both sides, so these are completion times, not time-to-first-token.

  • jev-model-tokengate - Security & Guardrails: An OpenAI-compatible proxy that sits between your LLM and your users. It evaluates each sliding window of tokens **while the response is still streaming** and cuts the stream **before** a violating token can reach the screen.

  • open-jev-approvals - Security & Guardrails: Binary approval gate for Codex and Claude Code — every intercepted tool call is reviewed by TypeSafe JEV and composed through a versioned local policy, with scoped authorization.

  • reflex - Security & Guardrails: A coding agent and personal assistant built on the Pi coding agent. Jev checks every tool call, turn and voice transcript, and code decides what happens next: allow, ask or block an action, which model tier to use, and whether a "done" was actually verified.

  • ego-jev-ultrafast - Security & Guardrails: Jev drives your Ego Lite browser: one typed-choice request per step. Single-file, zero-dependency port of browser-use/jev-ultrafast with multi-model benchmarks and extra guardrails. Unofficial.

  • jev-decisions - Security & Guardrails: Jev Decisions Plugin for Hermes (and other AI Agents): tool risk reviews, human approval recommendations, evidence checks, and a local decision journal.

  • jev-guard - Security & Guardrails: High-speed, cross-agent safety gate plugin for **Claude Code**, **Codex CLI**, and **Antigravity**.

  • jev-secret-detection - Security & Guardrails: Benchmark and tool evaluating how well TypeSafe Jev identifies real secret credentials in file snippets.

  • jevshield - Security & Guardrails: Sub-100ms security gate for AI agent tool calls, powered by TypeSafe's Jev (System-1) decision model. Single-request Choice/Noul/Score evaluation, dual-factor blocking matrix, calibrated-confidence routing, fail-closed parsing, zero-config local fallback. LangChain-ready.

  • oc-plugins - Security & Guardrails: The oc-auto-perms plugin in an OpenCode plugin collection uses Jev to check tool intent against natural-language rules.

  • actiongate-jev - Security & Guardrails: Open-source Jev tool-calling authorization gateway for AI agents: deterministic policy, exact-action single-use permits, MCP and HTTP enforcement.

  • antivirus - Security & Guardrails: A file scanner that sends extracted features to Jev for a verdict, a 0–4 severity score, and Noul indicators, then applies local quarantine or review rules.

  • claude-jev-plugin - Security & Guardrails: TypeSafe Jev semantic guardrails for Claude Code

  • dsh-jev-verify - Security & Guardrails: Jev (TypeSafe System One) decision tools + live verification benchmark for DeepSeek Harness: jev_decision (choice/score/noul) and jev_verify, honest by design.

  • grok-jev-guard - Security & Guardrails: `grok-jev-guard` sits immediately before a meaningful Grok Bot tool sequence. It receives a compact description of the pending operation and returns one explicit action:

  • jev-cvss - Security & Guardrails: Scripts that use Jev to select CVSS metrics from vulnerability descriptions, then compute v3.0, v3.1 or v4.0 scores in Python.

  • jev-pii-checker - Security & Guardrails: A CLI that sends text to TypeSafe Jev for PII category Nouls and a sensitivity Score, then locates spans with regex and segmentation.

  • jev-risk-check-provider - Security & Guardrails: x402check — LIVE payer-intent risk checks for x402 agent commerce: typed decisions (TypeSafe Jev), ES256-signed attestations, mainnet USDC settlement. did:web:x402check.xyz

  • opencode-jev-guard - Security & Guardrails: When FarHand is active, the agent's commands run on a remote host through the `farhand_remote_shell` MCP tool instead of `shell`. OpenCode's permission request for an MCP tool carries no arguments, so the plugin takes the command (and its `cwd`) from the `execute.before` hook, which OpenCode runs first.

  • claude-jev-warden - Security & Guardrails: Real-time quality gate and Art Director Warden for Claude Code powered by TypeSafe Jev 1.13 non-autoregressive decision model

  • guardrail-chatbot-jev - Security & Guardrails: It is a library, not a service. You call it, you get a verdict, and your code decides what to do. It runs in **Python and TypeScript**, both reading the same policy file, so the two sides of your stack cannot drift apart. Neither package has a third-party dependency.

  • jev-chrome-extension - Security & Guardrails: 1. Open any website. 2. Click the Jev icon. The side panel opens on **Drive**. 3. Type a goal, e.g. *Search Wikipedia for "espresso" and open the article*, and press **Run**.

  • jev-preflight - Security & Guardrails: A bounded Jev risk check for Claude Code: eight risk axes, one request, one optional reinspection.

  • jev-reasoning-navigator - Security & Guardrails: En lugar de depender de heurísticas matemáticas frágiles o distancias vectoriales locales de coseno, `JEV-Reasoning-Navigator` utiliza **TypeSafe AI (`typesafe-sdk`)** como motor único y autoritativo de decisión cognitiva:

  • jev-test - Security & Guardrails: Prototype: AI-assisted NZQA marking from rubric criteria alone, using TypeSafe Jev for guardrails, criterion scores and confidence-based triage. NOT ENDORSED BY NZQA - CONCEPT ONLY

  • momus-review - Security & Guardrails: Code review on rust and javascript (+languages soon) applications. Following lenses correctness, security, reliability, compatibility and testGap.

  • pkg-gate - Security & Guardrails: Pre-install security gate for npm lifecycle scripts using TypeSafe System One. Evaluates preinstall, install, and postinstall hooks across intent, threat severity, secret access, and remote execution to intercept supply-chain attacks before execution.

  • traffic-guard - Security & Guardrails: High-throughput traffic and attack defense gate for incoming HTTP traffic with zero required dependencies, wire-order header validation, and TypeSafe System One acceleration for bot mitigation, exploit detection, and risk scoring.

  • laya-browser-guard - Security & Guardrails: A passive, privacy-first Chrome security copilot that combines deterministic browser-visible checks with local Laya and official Jev typed decisions. It evaluates redacted evidence from scripts, resources, forms, headers, and DOM signals, then explains investigation priority without attacking the target.

  • Edward - Agent operations: one batched Jev Choice over the cross-turn coding-agent trajectory decides continue, pause, or escalate, with low-confidence verdicts routed to a human while deterministic code keeps dangerous-command blocking, budget caps, and an Ed25519-signed receipt chain.

  • taste-lint - Writing / UI: CLI that uses Jev probabilities on semantic taste checks to catch AI slop in UI, copy, and agent instructions before ship; measurable rules stay local and active findings can fail a run.

  • jev-harness agent: Multi type: cli - Developer tooling: System 1.5 quality gate and token optimizer for AI coding agents that triages test failures in < 500 µs to resolve missing dependencies without frontier LLMs, aborts circular doom loops, and modulates reasoning effort across Python, TypeScript, and Rust.

  • TryJevAI - Scheduling: public Jev playground uses a typed Choice with an explicit Unresolved option to distinguish a mentioned arrival time from an agreed meeting time, showing the returned probabilities and prompting for missing agreement before treating a time as settled.

  • Agent Chaperone - Agent safety: MCP proxy plus hooks that screen a tool call before it runs and a tool result before the agent reads it, with 45 test files behind it.

  • jev-proof - Creator sponsorship: verifies each sponsored ad segment in video subtitles against acceptance rules with one Jev Noul+Choice call per fact while deterministic code keeps the confidence gate; 90-sample calibration reports 0.922 gated accuracy and 0/15 injection flips.

  • jev-fidelity - Editorial QA: asks Jev per fact unit whether an edit preserved the original (preserved / equivalent / drift / lost) behind a 0.70 confidence gate in code, degrading to human review rather than pass; 55-sample calibration on real Wikipedia revision diffs reports 91/92 gated judgments correct and 0/20 injection flips.

  • approval-judge-bridge type: proxy - Agent safety: OpenAI-compatible /v1/chat/completions proxy that gates an agent's shell commands through a calibrated Jev Choice decision with fail-closed semantics.

  • Dub - Link safety: calls typesafe-ai/jev in malicious-link-check.ts before a short link is created, so the URL is gated by a typed verdict rather than a blocklist.

  • JevGate type: cli - Code review: CI and coding-agent gate that parses code locally and asks Jev Noul, Choice and Score questions about one function, file outline, candidate copy pair or test at a time, turns answers at 0.80 into review or consider findings with file and line, fails the build on review, and keeps undecided files as uncertain instead of clearing them.

  • dsh-jev-interceptor agent: DeepSeek Harness type: plugin - Coding agents: DeepSeek Harness plugin where a Jev Choice risk class plus Noul irreversibility, task-match, and injection checks gate every non-read-only tool call (deny confident high-risk, ask ambiguous, delegate the rest), Noul scope and reversibility questions auto-approve clearly-granted calls behind argument-evidence gating, and a per-message Score re-ranks what a referenced session keeps instead of oldest-first dropping — fail-closed to stock behavior, shadow mode with a /jev-stats command, 64 tests.

  • jevci type: cli - Quality gate: asks four typed questions about each change — three Score lenses and one Noul — and blocks a diff, commit message or doc set that falls below the resulting quality score, from the terminal, a pre-commit hook or a GitHub Action.

  • pi-subagent-jev agent: Pi type: plugin - Agent governance: when the Pi main agent dispatches a subagent, evaluates the task text against configurable rule sets in one typed Jev call (per-rule probability questions with below/above thresholds), blocks the dispatch with per-rule reasons on any hit, and fails open to allow on errors.

  • jev-lint agent: Multi - Software engineering: uses Jev Noul judgments and local thresholds to flag team-rule violations as Claude Code and Codex edit, helping agents fix them before code review with configurable rule packs and repository-specific rules.

  • jev-secret-guard agent: Claude Code - Agent security: Claude Code PreToolUse hook that blocks known key formats locally and sends unknown high-entropy strings to Jev only in masked form for a Noul on whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable; 6 of 6 secrets and 0 of 6 benign strings were blocked in its published calibration.

  • Perch agent: Multi type: cli - Code linting: semantic code linter that asks Jev about each method with its callers and callees in view, a Noul for whether it has a bug, a Choice for which kind and which line, and a Score for severity, plus language-filtered CWE Noul checks and custom rules written as sentences at repository, file or method level, failing CI on any answer over its floor.

  • semcheck type: cli - Code review: Go linter whose rules are plain-English questions such as "does this log call write personal data?", asking Jev one Noul for each piece of code a rule applies to and reporting it above the rule's threshold; its two shipped rules were right on 12 of 12 sampled findings in three open-source projects.

  • Skill Scanner - Agent security: Cisco's scanner hunts prompt injection and exfiltration in agent skills, and ships a System One analyzer as a deliberately advisory tier that cannot emit a finding or change a severity.

  • openclaw-jev-leakguard agent: OpenClaw type: plugin - Agent security: OpenClaw plugin that checks every outgoing agent message against where it is going, running local key-format and term rules and then five Jev Noul questions in one call (credential, where credentials are kept, client name, internal infrastructure, confidential business information) through OpenClaw's decisionModel, hosted Jev or a local Kev, and blocking, asking or holding it back by the channel's public, shared or private tier; with Jev it missed 0 of 56 synthetic leaks, 30 of which no regex or term list could see, with 4 false alarms on 57 ordinary messages at 223 ms p50.

Scoring & Ranking

Source file: categories/scoring-ranking.md

  • Clean Code Judge - Code quality: scores every file of a pull request on 31 boolean Clean Code smells plus function size and nesting, then hands the verdicts to a writing model for the review prose.

  • citation-verifier - Academic publishing: checks whether each cited paper actually supports the sentence citing it, with Claude locating the quote, Jev scoring the support, and a human making the final call.

  • jev-bfs - Search tooling: finds link paths between English Wikipedia articles by having Jev rank each page's outgoing links while Python controls the search.

  • Jev Search - Web search: uses Jev Noul judgments on result titles and snippets to rank Search1API results by relevance, with application code merging duplicate URLs and grouping lower-scoring matches separately.

  • pagegrade - Content quality: grades page sections for clarity, writing, and on-page SEO with Jev and returns per-section scores.

  • jev-scout - Developer tooling: sub-second zero-hallucination open-source repo and crate scout using TypeSafe Jev speculative fan-out scoring.

  • jev-seo - Zero-cost, agent-first SEO & Generative Engine Optimization (GEO) search radar CLI suite and MCP server powered by DuckDuckGo and TypeSafe Jev System One.

  • JevSlop - Writing quality: scores public note.com articles on eight Jev Score axes inside a single systemOne request and turns them into a 0-100 Slop Score in ordinary TypeScript.

  • SemanticSpace - Semantic mapping: places phrases in 2D by asking Jev how strongly each one relates to two chosen axis concepts and using those scores as coordinates.

  • Supercov - Code quality for coding agents: Jev answers twelve Noul properties per source file so the agent knows what to fix first.

  • jev.nvim - Developer tooling: Neovim plugin that splits the buffer into functions with Treesitter, scores each against a plain-language question with Jev, and ranks answers by probability in quickfix.

  • jev-reranker - Retrieval and RAG: uses Jev Noul judgments to assess retrieved documents for relevance and usefulness as answer evidence, then sorts results and optionally filters them using a configurable threshold.

  • jev-skip - Media: browser extension that reads the YouTube caption track and scores each segment's sponsor probability on the seek bar before the intro ends, reporting 77% of SponsorBlock's sponsor seconds caught over 23 videos at $0.0008 a video.

  • Refix - Growth: AI that helps your product grow faster on autopilot by running product experiments, SEO, content, and ads.

  • jevsearch - Site search: shadcn/ui ⌘K search block that streams local keyword hits, then sends the top 20 to Jev in one request (a Noul per candidate, a Choice for the best page, a Noul for whether any page answers) to re-order or drop hits, reporting Hit@1 of 83% versus 41% for keyword search alone on 41 labelled queries over the TypeSafe docs (author's benchmark).

  • JevPDF - Document search: browser PDF viewer that extracts each page's lines with pdf.js, asks Jev one Noul per line ("does this line answer the query?") in batches of up to 16 lines sharing the page text as state, and highlights lines ranked by probability as each page returns; only text reaches Jev, through a key-holding proxy.

  • slop-grader - Content quality: CLI tool that grades text files against custom rulesets for AI slop, grammar, and technical doc quality using Jev scores and line-level flags, then guides an AI agent to auto-fix violations.

  • jevseo - SEO and AI-answer visibility: a deterministic crawler extracts every page of a business site, Jev answers a narrow typed Choice/Score/Noul question set per page, and application code turns those probabilities into ranked findings under three confidence bands with the grey zone routed to a needs-a-human pile rather than acted on; it runs locally on one port with no API key through the keyless Zen tier, publishes no search-volume numbers at all because Jev carries no index or volume data, and its source is UNLICENSED (all rights reserved) - unrelated to the other jev-seo entry above.

  • jev-ai-detector type: extension - Writing analysis: Chrome extension which gives readers an instant, uncertainty-aware signal for how strongly selected webpage text resembles AI-generated writing, using Jev inline in Chrome without interrupting reading.

  • Tweet Radar - Social reading: uses Jev Noul to score already-loaded X posts against a reader's goal and profile, then pairwise Choice judgments to rank eligible matches and surface up to three for review.

  • nlgrep - Developer tooling: uses Jev Noul judgments to find code, docs, logs, and text satisfying natural-language conditions, with a configurable probability threshold and ranked file results linked to source lines.

  • hippo-memory - Agent memory: a biologically-inspired memory store whose optional Jev reranker lifts recall R@1 from 0.41 to 0.62 on a private 300-query developer store.

  • MemSearch Jev reranking - Coding-agent memory: an optional Jev reranker asks Noul questions about retrieved Markdown chunks and sorts them by relevance to the query, with bilingual evaluation results.

  • Oko - Developer tooling: local code search for coding agents that shortlists function-level chunks with ripgrep and BM25, asks Jev a Noul relevance question per chunk across three parallel requests, and returns the accepted ones as excerpts through MCP; the cutoff and excerpt selection live in code.

  • grokbot-jev-jobs - Job search: a daily Vercel cron that scores public job postings against one resume with Jev through the Vercel AI Gateway, so only the plausible matches surface.

  • jeff type: cli - Developer tooling: Go CLI whose rank command asks one Jev Score per item per weighted dimension of a YAML spec in a single request and sums weight times score in code to order the items, with noul, choice and score commands that turn a threshold into exit code 10 for shell scripts and CI.

  • OpenViking - Reranking: Volcengine's agent context database ships a Jev rerank client that scores each candidate document with jev-latest against api.typesafe.ai and treats the returned probability as relevance, because TypeSafe exposes no native rerank endpoint.

  • Jev-Code-Reviewer agent: Multi type: cli - Code review: asks Jev for a priority score per changed unit and returns a priorityGap that a local uncertainty policy turns into the order a human should read the hunks in, while OpenAI explains the ones that surface.

  • WorldMonitor - Geopolitical intelligence: real-time global intelligence dashboard using TypeSafe Jev questions to score news headline severity into 5 threat levels and categorize events across 14 conflict, cyber, and infrastructure domains.

Agent Decisions

Source file: categories/agent-decisions.md

  • jev-social - Social media research: uses a Jev Choice at each step to select a concrete socai CLI operation and observed post or profile target on Instagram, TikTok, or LinkedIn, rejecting malformed or low-confidence decisions before execution.

  • Jev Ultrafast - Browser automation: browser-use's ultrafast agent where Jev decides each next action and element to click, calling a language model only when text must be typed.

  • jev-agent-browser - Browser agents: a parent agent delegates bounded tasks to a Jev loop that selects typed browser actions, validates them through agent-browser, and escalates ambiguity or stuck states back to the parent.

  • pi-typesafe-jev - Coding agents: exposes System One judgments as five Pi tools so a model makes narrow semantic judgments while code and users keep control of thresholds, weights, and actions.

  • jev-judgment - Coding agents: agent skill that sends closed coding-agent judgments to Jev so verdicts stay typed, cheap, and comparable across runs.

  • limpet - Coding agents: Stop hook that keeps an agent from finishing too early by judging plain-language completion rules with Jev.

  • robo-harness - Robotics: SO-101 arm workbench where a Jev decision runner picks bounded joint steps from typed candidate actions under a spend budget.

  • dsh-auto-mode - Coding agents: DeepSeek Harness permission preset whose end-prompt step has Jev answer the open questions an agent leaves in its final message, steering them back only when a choice clears 0.6 confidence and an autonomy-safety Noul clears 0.5, and returning the turn to the human otherwise.

  • augustus - Coding agents: agent skill that maps Choice, Score, and Noul onto classical methods so an agent can place typed judgment in software, with a composition algebra, question-design diagnosis, and a validation gate that requires a falsifying experiment.

  • yoshi - Context management: proxy for Claude Code and Codex where Jev judges which conversation history is still needed before pruning.

  • pi-jev (TheoOliveira) - Coding agents: semantic tool routing and typed System One decisions for the Pi coding agent.

  • pi-quiet-ask - Coding agents: gives the Pi agent a quiet Jev decision layer for judgments it would otherwise hand to a chat model.

  • fastbrowse - Browser agents: Jev picks each action from what is on the page while an LLM reads and plans.

  • super-jev - Decision harness: turns a Jev answer into a bounded action instead of leaving the caller to interpret it.

  • jev-superpowers - Systematic software development framework for AI coding agents upgraded with TypeSafe Jev System One typed decisions, zero-hallucination package vetting, and completion gates.

  • Jev Browser - Browser automation: drives a browser with Jev deciding each step, pitched as fast and very cheap next to LLM-driven browsing.

  • pi-fast-jev-compaction - Context management: Pi extension that keeps conversation text verbatim while pruning stale tool history with Jev, falling back to Pi's own summarization only when pruning cannot free enough room.

  • Atomic - Coding agent runtime: ships a first-class Jev structured-output provider so an agent's decisions come back typed, through the same decision resolver as its other providers.

  • fast-jev-compaction - Context management: Claude Code plugin that replaces the compaction summary with Jev decisions, scoring every tool call and result for whether it is still needed instead of summarizing the session.

  • fast-dev-compaction - Context management: Codex port of the Jev-guided compaction idea, restoring context verbatim around a session compaction rather than summarizing it.

  • public-browser - Browser control: lets Claude Code and Cursor drive a real Chrome profile, with a Jev loop deciding the actions, reporting roughly 30% fewer tokens and 25% lower cost.

  • pi-typesafe-router - Coding agents: routes Pi's work through typed Jev decisions.

  • wakegate - Long-running agents: before a sleeping agent's LLM is resumed on a timer or incoming event, Jev answers a Choice (wake, not yet, unrelated) against the agent's own sleep note, and code skips the wakeup only when wake is below 0.2 while always waking on user messages, bare timers, a skip limit, errors, and timeouts; one run passed 21 of 21 hand-written scenarios, which the README calls a smoke test rather than a benchmark.

  • BrowserClaw - Browser automation: Zero-lock, session-preserving Chrome MCP server that couples a local Jev System One semantic micro-loop (chrome_act_toward_goal) with an 85%+ pruned DOM tree (Shadow DOM & iframe pierced), dispatching native CDP events (isTrusted: true) on active logged-in sessions without focus theft.

  • jev-canvas - Multimodal UI: draw on a tldraw canvas by voice while pointing a webcam-tracked finger; on every partial transcript Jev answers eight typed questions (is it a command, is the sentence complete, action, shape, colour, target, place, size) and plain code gates them with thresholds, in English and Ukrainian, 300–550 ms per decision.

  • jev-belay - Coding agents: Claude Code Stop hook that reads the transcript for evidence and spends one four-question Jev call only when files changed with no passing check since, failing open on any error.

  • Jev for Chrome - Browser automation: unofficial Chrome extension port of Jev Ultrafast where a Jev Choice picks the operation and DOM element each step and two Noul checks (goal reached, stuck) veto a premature DONE or BLOCKED, with a small text model used only when text must be typed.

  • Jevonian - Coding agents: local OpenAI / Anthropic / Responses-compatible proxy where one Jev call answers both the model route and the thinking level for jevonian/auto from session state, quota health, candidate capabilities, and cache-switch penalties; deterministic code filters candidates and owns every threshold first, a pinned model or explicit jevonian/<route> skips Jev entirely, and each decision lands in a local ledger with the serving model, reason, real token usage, and estimated cost.

  • jev-pruner - Context management: Claude Code plugin that trims long Bash output with Jev before the model ever sees it, keeping terminal noise out of the window.

  • jev-desktop - Computer use: supplies Jev action selection inside Codex Computer Use, choosing among desktop actions rather than asking a language model at every step.

  • jev-browser-bridge - Browser agents: plugs any CDP browser into a Jev loop, where a Jev Choice picks the operation and its target element each step from candidates read off the DOM rather than the layout, so the same agent runs on Chrome and on engines that never draw a page (Moli, Lightpanda, Kitesurf), passing at least 90% of runs on each of fourteen browsers tested.

  • Sedum - Browser end-to-end testing: in goal mode each turn asks one Jev Choice for the next operation (click, type, done or blocked) plus a speculative target among the page's offered elements, capped at 24 requests, 18 actions and 120 s, and the test passes only when an independent verify claim clears two Nouls (holds ≥ 0.75, contradicted flagged at ≥ 0.5), since the planner's done is never a verdict; authored-step tests reuse the same Choice to resolve each plain-English step, and with your own API key a 20-person team's PR suite costs $38–$91 a month vs $4,875 on a per-step AI platform.

  • killmyidea - Decision Tools: A startup-idea evaluation demo assigning KILL, FIX or SHIP labels from Jev scores.

  • jevify - Decision Tools: An Agent Skill for finding suitable Jev decision points and designing questions and comparison experiments.

  • hermes-jev - Decision Tools: An asynchronous Jev companion for Hermes Agent covering relevance, completion, recovery and optional admission decisions.

  • claude-jev - Decision Tools: Claude Code plugin: Jev for rule checks, verbatim compaction, and prompt routing

  • wechat-jev-assistant - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-chat-windows-deepseek-jev - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • Jev-chat-assistant - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-skill-router - Decision Tools: Claude Code plugin: asks TypeSafe Jev which installed skill fits each prompt and logs the answer (shadow-first). A working reference for the skill-suggestion cookbook on Claude Code — the README records why it is unlikely to help a strong model as a router.

  • jev-apply - Decision Tools: In Codex, Claude Code, or another CLI agent:

  • jev-bot - Decision Tools: Self-hosted Jev decision workbench and Feishu bot: automatic choices, probabilities, and experimental word/character writing.

  • dsh-jev - Decision Tools: DSH bundle that registers jev_ask for TypeSafe Jev noul, choice, and score answers.

  • jev-chat-windows-laya - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-demos - Decision Tools: Every demo lives in its own folder with its own README, dependencies and instructions. Most run without an API key in a clearly labelled `SIMULATED` mode; put `TYPESAFE_API_KEY=...` in the demo folder's `.env` for live results.

  • Jev-in-the-Loop - Decision Tools: Researching how Jev can accelerate tasks that rely on LLM decision-making.

  • jev-laya-benchmark - Decision Tools: Speed and accuracy benchmark: TypeSafe's Jev API vs the local Laya MLX typed-decision model on synthetic tasks

  • jev-predict-skill - Decision Tools: An Agent skill recipe that predicts another skill’s closed-set outcome from its rules and evidence.

  • astra-jev-harness - Decision Tools: The default `batch` policy retains an entire batch when any judgment is uncertain. The experimental `select --policy per-file` retains uncertain/unjudged files while omitting confidently irrelevant siblings. Dependencies and the global no-match fallback still apply. Compare before changing policy; fewer bytes alone do not establish correctness.

  • jev-codex-router-skill - Decision Tools: Portable Codex Skill for Jev model and reasoning-effort routing, with safe installation and Chinese usage guides

  • jev-playground - Decision Tools: A web playground for entering state and decision questions, then inspecting Jev answers and probability distributions.

  • fake-real-jev - Decision Tools: See link entry, the live scan timer, JEV's evidence-checking role, a saved REAL example, a saved FAKE example, and the linked sources. The live documentation scan shown ended without a verdict; its credit was returned. The coffee reports are clearly labeled saved examples, with their original analysis times.

  • jev_projects - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-crawlers - Decision Tools: Jev learns your repo's decision norms, then adversarially judges past decisions against them. Unix-style primitives (seed, expand, judge, verify, report, norms) with per-node typed judgments from typesafe-ai/jev.

  • jev-linter-action - Decision Tools: `glob` accepts one pattern or a newline-separated list. All matched files are reviewed together by default; `per-file: true` reviews each file independently. Missing inputs, malformed questions and ambiguous combinations fail before calls.

  • jev-no-enem - Decision Tools: Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.

  • jev-triage - Decision Tools: Millisecond-class test-failure triage for coding agents: RETRY / FIX_CODE / FIX_ENV, powered by TypeSafe Jev.

  • Jevatar - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jevchat - Decision Tools: A chat-style Jev demo whose answers are selected from predefined or custom options rather than generated prose.

  • JevCode - Decision Tools: JevCode - Jev can code. We want to dogfood JevCode

  • typesafe-jev-ruby - Decision Tools: Ruby client for Jev, TypeSafe's System One model: typed questions, probabilistic answers. Zero runtime dependencies.

  • xjevboost - Decision Tools: Add as much tabular data as you want to Jev models using adaptive ensembles that learn to query only the rows and columns needed.

  • turing-jail - Decision Tools: Interactive three-level AI interrogation game powered by TypeSafe Jev; write responses and pass plea, logic, and paradox verdicts to earn release.

  • Hermes JIT Context OS - Coding agents: uses Jev as a sub-millisecond System 1 Epistemic Gate and Domain Router to score AST relevance, test proofs, and tool targets, cutting autonomous agent turns by 31.3% and blind file exploration by 52.6% on SWE-bench with fail-open circuit-breaker resilience.

  • jev-agent-skill agent: Multi type: plugin - Developer tooling: Claude Code/ZCode skill that offloads classify/route, batch-screen, score, and compliance-check judgments to Jev via OpenCode Zen's free tier, bundling a zero-dependency jev.py caller (transient-500 retry, WAF-safe UA, GBK-pipe-safe stdin) and a production Taobao-shop comment-triage pipeline that keeps raw items out of the agent context.

  • Yappy - Computer use: macOS voice agent that asks Jev one Choice per step (operation and target control) over the front window's accessibility table, executes only validated high-confidence answers, and escalates to a full LLM agent on low confidence, no-effect actions, or unknown field values; author-reported 275–690 ms per decision.

  • JevLoop (parkavenue9639) - Agent runtimes: a Python runtime where Jev Choice decisions select tools and targets, uncertain decisions escalate to an LLM, and a shared guarded kernel supports isolated Docker workspaces and paired LLM-only comparisons.

  • GUI JEV Harness - Computer use: recursive screenshot grounding where Jev returns a Choice over grid-tile candidates at each level, and local probability and margin gates decide whether to descend or refuse, emitting only a raster point and bounding box and never clicking.

  • Visual-JEV - Multimodal models: Jev-style model built on Qwen3.5-4B that takes images directly, without first converting them to text.

  • DeepSearcher stopping-policy experiment - Agentic search: a standalone evaluation uses Jev Noul judgments on accumulated evidence to decide whether to stop or continue within a search-round budget, comparing stopping behavior, evidence recall, and decision cost.

  • jev-chat - Messaging: an Android accessibility service reads the conversation in WeChat, QQ, X, or Feishu, asks Jev Choice over candidate replies, and fills the draft box while sending stays manual; a Windows port does the same from offline OCR of the WeChat window.

  • SkillRanker agent: Claude Code type: cli - Coding agents: standalone Rust CLI that uses Jev to rank candidate skills against live session context, advising the next step through a Claude Code UserPromptSubmit hook.

  • AutoGPT - Autonomous agents: open-source autonomous agent platform featuring first-class TypeSafe Jev decision blocks for typed routing, filtering, scoring, and confidence-gated next-action dispatching.

  • oh-my-claudecode agent: Claude Code type: plugin - Coding agents: multi-agent team orchestration for Claude Code featuring opt-in Jev hooks for sub-millisecond judgment points, decision caching, and per-point egress controls.

  • jcode - Agent runtimes: RAM-efficient autonomous agent harness implemented in Rust with native TypeSafe Jev typed decision transport for memory pruning, browser navigation, and voice interaction routing.

  • opencode-jev-compaction - Replaces OpenCode compaction summaries with Jev keep/drop judgments that prune stale tool calls while preserving everything kept verbatim.

  • jev-auto-approve agent: Claude Code - Coding agents: Claude Code PreToolUse hook that asks Jev a Noul on whether a shell command is strictly read-only, auto-approving at 0.95 and otherwise falling back to the normal permission prompt without ever denying, while a local hard-no list and injection filter keep risky commands from reaching Jev; 0 of 8 state-changing commands were approved in its published calibration.

  • laya-browser-agent - Browser agent: derives each step from a Jev-shaped model — Laya through MLX or PyTorch, any duck-typed backend, or an arbitrary System One HTTP endpoint.

  • openclaw-jev-trigger agent: OpenClaw type: plugin - Agent automation: OpenClaw plugin and CLI that turn a plain-language --when / --not-when condition into a scheduled trigger script, asking Jev one Noul per tick through OpenClaw's decisionModel and waking the conversation model only when the condition becomes true at 0.7 or above; on 76 synthetic watcher ticks Jev was right on 75 with 0 false wake-ups at 231 ms p50 and about $0.000016 per check, against 87% for first-try JavaScript rules.

  • mu - Coding agents: a pi-based coding agent and desktop app that asks Jev Noul, Choice and Score questions at 38 decision points inside its loop, such as which chunks of a long tool output enter the context, whether a rule-flagged command was actually asked for, whether a fetched page or MCP result carries instructions aimed at the model, and whether a "done" was verified; each point acts on its verdicts by default, can be switched to shadow or off, and writes every verdict to a local ledger, and in the repository's replay benchmark goal-aware test-log selection by Jev cut 40-46% of the log without losing a required line.

Data Labeling & Curation

Source file: categories/data-labeling-curation.md

  • jev-curate - Dataset engineering: sifts synthetic JSONL and Parquet rows using Jev Noul checks and calibrated confidence scores, streaming passed records and rejections straight to disk.

  • typeful-triage - Open-source maintenance: multiplayer triage dashboard where Jev answers a fixed set of typed questions per issue — kind, severity, urgency, duplicate, and next step — and every human correction is kept and shown back to the model on later runs.

  • JevSpan - Information extraction: zero-shot named entity recognition that splits text at punctuation, asks Jev one Choice over every candidate window per entity type, verifies each nominee with a second Choice (the type, none, mixed or partial) and settles its boundary with a third, averaging 73.7 strict F1 across 12 Chinese and English NER benchmarks against 72.1 for direct extraction with Qwen3.8-27B.

  • GroundingJev - Visual annotation: a Jev-inspired Qwen3.5-0.8B model that maps an image and referring expression to four bounding-box coordinates in one forward pass, reporting an 8.61× inference speedup over its autoregressive base model.

Evaluation & Benchmarking

Source file: categories/evaluation-benchmarking.md

  • Jev Playground - Model evaluation: benchmarks Jev against Luna, Haiku, and Gemini at choosing validated legal moves in explicit-state games, scoring decision quality and consistency across a sequence of moves.

  • Jev vs Mistral and Gemini for event validation - Event discovery: head-to-head test of Jev against Mistral Small and Gemini Flash-Lite at validating local event listings.

  • jev-research-eval - Research automation: reproducible eval harness plus field note for Jev Ultrafast research-browser tasks, with QC'd cases, a suite runner, and a report generator.

  • Jev judge call vs dimension scores - Model evaluation: tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

  • Jev Pong - Model comparison: Pong where the ball advances one step per model decision, putting Jev head-to-head with LLMs through Vercel AI Gateway.

  • Jev reranking is not a free win - Search reranking: a measured run over 33,047 catalog entries, 164 real queries, and 9,831 graded pairs reports that Jev reranking alone did not beat vector retrieval.

  • An early-access test of TypeSafe's Jev - Independent trial: measures calibrated judgments on early-access Jev and reports the resulting cost per decision.

  • jevcal - Model evaluation: fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, reports how much traffic still has to escalate to an LLM, and fails CI when a model update breaks the locked thresholds.

  • WindTunnel - Browser-agent benchmark: measures WebMCP against other browser-agent interfaces, with Jev appearing as one of the compared configurations.

  • jev-eval - Third-party check: compares Jev against GPT-4o-mini and Claude Sonnet 4.5 under identical conditions on the same judgment task.

  • minutes - Meeting notes: local-first transcription app whose live voice path runs its evaluations through Jev.

  • jev-orderby-bench - Model evaluation: measures whether a SQL ORDER BY over a Jev probability is defensible (pairwise inversion, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a pre-registered gate that jev-1.13.0 passes on 20 Newsgroups topics and fails four of six conditions on Amazon ESCI product relevance, and shows a DuckDB extension's default 40-row batching fails the ranking gate that one row per request passes.

  • jev-ood-calibration - Model evaluation: independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).

  • Odin R&D: Jev vs open-weight Laya - Model evaluation: runs the same 48 hand-written choice/score/noul rows through Jev's hosted /api/v1/systemone and a pinned open-weight Laya on a Mac (MLX, checked row-by-row against Laya's own reference code) under pre-registered refutation criteria, publishing the raw results record, 46/48 vs 39/48 accuracy, and a reproduce command.

  • latitude-llm - Evaluation & Observability: Latitude includes an optional Jev preclassifier for conversation checks and their selection records.

  • jev-review - Evaluation & Observability: A local MCP code-quality reviewer returning structured scores to coding Agents.

  • taskuary - Evaluation & Observability: An optional Jev judgment module in Taskuary for checking user-defined conditions on task state.

  • Canny - Evaluation & Observability: Keeps an execution ledger for Claude Code and Codex CLI to check for passing validation after edits.

  • jev-playground - Evaluation & Observability: A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.

  • typesafe-ai-benchmark - Evaluation & Observability: Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.

  • goodwatch-monorepo - Evaluation & Observability: A film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.

  • jev-benchmarks - Evaluation & Observability: A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.

  • jev-rerank-bench - Evaluation & Observability: Compares Jev, dedicated rerankers and chat models on the same retrieved passages.

  • jev-benchmark - Evaluation & Observability: Benchmarks Jev on chess moves and identifying which game NPC a player addresses.

  • jev-lm - Evaluation & Observability: A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.

  • jev-chat - Evaluation & Observability: A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.

  • jev-frontend-qa - Evaluation & Observability: Frontend QA that uses Jev to choose browser actions and checks contracts through DOM, HTTP and database evidence.

  • jev-behavior-study - Evaluation & Observability: An independent Jev 1.13.0 behavior study recording successes and failures across question framing, input conditions and games.

  • jev-exploration - Evaluation & Observability: A research repository tracking Jev claims and limitations, with calibration experiments and runnable examples.

  • jev-gomoku - Evaluation & Observability: A nine-board, 15×15 Gomoku workbench comparing how two Jev players respond to different input representations.

  • jev-agent-failure-benchmark - Evaluation & Observability: A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.

  • ask-jev - Evaluation & Observability: A Windows PowerShell tool for auditing recorded Codex execution evidence with :jev.

  • jev-synergy-screening - Evaluation & Observability: A Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.

  • hermes-jev-north-star - Evaluation & Observability: A Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.

  • jev-calibration-audit - Evaluation & Observability: Audits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.

  • jev-demos - Evaluation & Observability: Maze experiments comparing Jev single-step choices with multi-step lookahead.

  • jev-eval - Evaluation & Observability: Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.

  • jev-flash-review - Evaluation & Observability: An MCP review engine that evaluates Agent-supplied diffs against explicit rules.

  • foreman-jev - Evaluation & Observability: An experimental Jev supervisor for Codex workers with programmer-selected acceptance commands.

  • ASSAY-001 - Independent pre-registered check of Jev calibration and type safety on Banking77 / CLINC150, with a split verdict and full logs, written up at donttrustme.ai.

  • BTK audit studies - Content & growth: Jev striking-distance triage ranks SEO fixes and drives study pages; 1,204 pages judged per run, 4,816 judgments in under 3 minutes, $0.0048 per 12-query batch.

  • Can Jev Be a Better Agent Evaluator? - Agent evaluation: LangChain compares Jev against LLM judges on accuracy, repeatability, latency and cost, concluding Jev is the cheaper and more consistent judge for online evals.

  • jev-acento - Language evaluation: pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish state costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing instructions in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.

  • Jevals.com - Model evaluation: independent leaderboard that asks Jev and six LLMs the same Noul, Choice and Score questions and grades every answer against human labels (PubMedQA, Banking77, HelpSteer2; 300 items × 5 runs each), finding Jev tied for first on PubMedQA yes/no at 1/28 of the top LLM's price, tied for second on Banking77 and no model beating the label base rates on HelpSteer2, with every per-decision probability published as CC BY 4.0 data.

  • Jev IDS - Network security: an Intrusion Detection System prototype that takes the metadata of a network flow and returns a verdict on whether it is an attack plus its threat category with probabilities, and on the NSL-KDD benchmark was 4.8× faster and 3.8× cheaper than a state-of-the-art LLM (GPT-5.6 Luna) while raising 15× fewer false alarms than a Random Forest model.

  • jev-test - Model benchmarking: reproducible test harness evaluating TypeSafe Jev Noul, Choice, and Score decisions via OpenRouter's Decisions API, comparing latency and accuracy against LLM prompt-and-parse baselines.

  • Jev vs Fable on 520 real social posts - Social media: a scheduler's pre-publish check asks Jev four Noul questions per caption (spam, clear opening, stands alone, promotional) as advisory signals, never a gate; on 100 posts labelled blind by Fable the two agreed 94/100 on promotion and 85/100 at a 0.65 spam threshold (Jev the stricter one 12 times to 3), and scoring all 520 posts cost $0.011 at a 341 ms median.

  • Jev Does Not Play Dice - Model evaluation: asks Jev a Choice over the six faces of a hidden fair die 400 times; Jev selects face 1 on all 400 trials with 82.9% mean reported probability and 19.0% accuracy, then tests whether stated probabilities survive in synthetic forecast documents, where a 30% shortage risk comes back as 5.3% via Choice and 26.7% via Noul; raw responses and analysis code on GitHub.

  • DecisionBench - Model evaluation: scores Jev Noul, Choice, and Score answers on pinned document-grounded tasks, counting malformed probability distributions as misses so model comparisons remain reproducible.

  • jev-regress-bench - Agent regression testing: after a config edit, one Choice (same / fact_differs / action_differs / specificity_differs) decides which of an agent's approved answers changed meaning rather than wording, and on 109 before/after pairs whose ground truth is derived from what each config rule does to the answer, Jev catches all 19 real changes with 13 false alarms against 33 for a markers-then-embeddings-then-LLM stack and 19 for the LLM judge alone.

  • jev-fanout-bench - Model billing: compares batched with one-question-per-call requests across 2,976 calls to jev-1.13-20260917 through OpenRouter's TypeSafe-compatible /systemone endpoint, reporting about 261 fixed input tokens per request, zero spread in the implied per-request cost across question counts, batched-vs-single answer differences comparable to repeat-request noise, and median input-token savings of 76–86% at eight questions.

  • SystemOneHarness type: cli - Model evaluation: execution harness and dual-loop test framework that compiles goals, browser environments, and MCP servers into bounded System One reflexes, evaluating Jev against deterministic baselines.

  • judgekit type: cli - Model evaluation: runs declarative YAML judgment tasks natively on Jev Choice/Score/Noul or any OpenAI-compatible backend (with a free rules fallback), gates low confidence at 0.7 (caught 3/3 misjudgments at 9% escalation, n=130), and publishes Chinese-scenario cost-accuracy numbers — 97.7% @ ¥0.105/1k decisions and 60.0% → 68.3% on a frozen 120-item human-labeled spam set at τ=0.10.

  • Convex Decision Evals - Model evaluation: asks Jev a Choice on 108 verified four-option questions about the Convex backend platform (no docs or tools in the prompt, each asked 3 times with shuffled options, random guessing 25%) alongside 14 LLMs, where jev-1.13 scores 84.6% at a 199 ms median and $0.0088 per full run against 98.0% at 2.12 s and $1.59 for the top model, with every answer, probability and raw request/response in a public explorer and the runner in get-convex/convex-evals.

  • jev-medhallu-benchmark - Medical AI: pre-registered test of Jev as a hallucination check on Stanford MedHELM's MedHallu (1,000 test items), asking one Noul on whether an answer misrepresents its PubMed abstract; Jev scored 92.9% against 92.4–95.1% for four fast LLMs at a 204 ms median and USD 0.03 per 1,000 checks, and letting Jev settle the 37% of items where it was at least 90% sure kept each LLM's accuracy with 37% fewer LLM calls.

  • zh-decision-bench - Benchmarking: first Chinese-language calibration benchmark for Jev-class decision models (378 items / 5 models incl. NeoHorse-Jev-4B; accuracy, ECE, option-order and zh-CN/zh-TW robustness; CC BY 4.0 dataset on Hugging Face).

  • jev vs. open alternatives - Document pipelines: compares Jev against open models and specialised tools on five chores — language detection, orientation, RVL-CDIP classification, bundle splitting and parse-tier triage — asked as Choice(2) up to Choice(16).

  • S1MB - Decision-model evaluation: compares Jev and open decision models across 137 English Choice, Noul, and Score benchmarks, including six synthetic generalization probes, with public evaluation data, recorded results, and an interactive leaderboard.

Calibration & Research

Source file: categories/calibration-research.md

  • decider - Open models: reproduces the System One shape with a Qwen3.5-2B fine-tune that emits typed decisions with calibrated probabilities in one pass.

  • openjev - Open research: independent local preview that answers bilingual probability questions from context, questions, and candidate answers, inspired by TypeSafe Jev.

  • Parallel Constrained Decoding (Qwen2.5-1B-RLCD) - Open research: RLCD-trained Qwen2.5-1B demo exploring open-source parallel constrained decoding as an alternative to Jev.

  • NanoJev - Open replica: a 0.6B parallel decision model that returns full probability distributions with no output-token decoding, shipped with its training pipeline, weights, and dataset.

  • open-alternative-jev - Open alternative: runs a Jev-shaped decision model locally on your own GPU.

  • mini-jev - Local reproduction: implements Jev's typed-decision interface on top of a local LLM.

  • Laya - Open alternative: non-autoregressive decision model that answers choice, score, and noul questions with RLCD-trained calibrated probabilities in a single ~35 ms forward pass, published on PyPI and Hugging Face.

  • Jev-compatible public API - Open research: a public Jev-shaped API backed by an open Qwen3.6-35B-A3B model so anyone can try the typed-decision interface.

  • kev - Trainable replica: a tiny Jev-like model on top of Qwen2.5-0.5B that trains and runs on a MacBook, shipped with its own research runs and evaluation scripts.

  • jevinci - Creative experiment: paints images by having Jev predict every pixel's colour in parallel, with predicted confidence deciding how wide each stroke is drawn.

  • jev-local - Local reproduction: Jev-compatible POST /v1/systemone server answering typed Choice/Score/Noul questions with confidence from open weights, verified as an official-SDK drop-in with temperature-fit calibration (set3 n=1316, 0.83 overall).

  • LitJev - Local reproduction: a reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.

  • ruling - Local reproduction: Jev-compatible POST /v1/systemone server that reads typed Choice/Score/Noul answers from the logits of any MLX checkpoint or OpenAI-compatible endpoint with no training, works as a drop-in for the official SDK, and replays Jev's published answers on 256 public judgments (231 vs Jev's 238, McNemar p = 0.21).

  • CUA-S1-FORMS - Specialist decision model: a 706,048-parameter, 2.8 MB jev-like option scorer that rates FILL / CHECK / CLICK / SKIP for each form field in one parallel pass, reporting 99.7% on its own form-filling eval against Jev's 83.6% - a specialist on home turf rather than a general win.

  • jevlike - Training library: build a small model that chooses among a changing list of text options and returns one probability per option in a single pass - the base CUA-S1-FORMS was built on.

  • jevbetter - Improved scorer: a stronger one-pass scorer over a variable list of text options, using a hashed n-gram encoder, rival-aware attention, and gated heads.

  • jevlike-esp32 - Edge deployment: exports a jevlike scorer as ESP32 firmware with a C scorer and a host-side check, putting one-pass decisions on a microcontroller.

  • von - Open alternative: a 395M non-autoregressive System One model that answers typed questions with calibrated probabilities in under 15 ms, positioned as a local drop-in replacement for Jev.

  • RSI-Jev - Trainable replica: an open 2B model answering Noul, Choice and Score on the same POST /v1/systemone wire format in one forward pass, over text or, since v4.0-VL, up to four images, researched and trained by a recursively self-improving AutoScientists loop that publishes every experiment it ran — 86.6 AUROC zero-shot on 2,162 VisA inspection photos (Jev-Omni 81.1, Gemma 4 12B 82.9) and 80.3% on five held-out image benchmarks.

  • Open Medical Jev - Medical evaluation: two frozen local readers answer one Noul-style yes/no probability per exam option, a fit-free router auto-releases items above the combined-confidence gate and escalates the rest, and a split-conformal candidate set bounds the error - landing within 2 points of hosted Jev on three 600-item national licensing exams with no fine-tuning, no distillation and no corpus.

  • OneJev - Open models: multimodal System One model in four sizes (0.8B to 27B); typed questions about a screenshot, photo, video or text get a calibrated probability for every option in one forward pass.

  • TetraJev - General decisions: two frozen open-weight readers give four readings per item (letter + per-candidate yes/no), fused fit-free and routed by agreement into auto-release or human review; evaluated across eight decision suites and a RAG reranking pass with published coverage–accuracy curves; no training, and it does not call the TypeSafe API.

  • Jebadiah - Open replica: Apache-2.0 decision models (27B, 9B, 4B on Qwen bases; bf16, GGUF and MLX) that answer Choice, Noul and Score questions with a probability for every option from one forward pass, and run anywhere: a standalone server with Jev's /v1/systemone wire and a playground, a llama.cpp script for the GGUF builds, or AINode (open-source local AI platform).

  • WebJev - Specialist decision model: Apache-2.0 open-weight Qwen3.5-35B-A3B fine-tune for browser agents, served by vLLM behind the same POST /v1/systemone and /api/alpha/decisions routes so a Jev client switches by changing only the base URL and key; inside the unchanged jev-ultrafast agent it completes 38.52% of 125 hand-picked real-website tasks graded by deterministic verifiers, against 16.67% for Jev 1.13.

  • JevForge - Open research: an end-to-end stack for auditable data construction, Qwen3.5-0.8B training, fixed Mind2Web and OOD evaluation, local serving, and a preliminary RLCD baseline.

  • Luce - Open recipe: describe the decision task in a sentence, an LLM teacher writes the training data, a LoRA + decision head on Qwen3-4B-Base answers choice/score/boolean questions with calibrated probabilities in one forward pass; trains on a 12 GB card, and reports accuracy and ECE next to Jev on identical test items (rule-generated tickets 91.1 vs 75.1, phishing 97.4 vs 62.6, GitHub issue priority 41.1 vs 37.5) with a browser replay demo that needs no GPU.

  • poorjev - Local reproduction: implements Jev's typed Choice/Score/Noul interface on commodity zero-shot NLI models and makes the confidence honest with temperature scaling and conformal abstention, shipping a reproducible calibration eval (ECE 0.170 to 0.071, cross-validated) that runs offline with no API key.

  • openJev-verdict-2.0 - Open decision engine: a calibrated 151M non-autoregressive model that reports beating both TypeSafe Jev and Laya on typed-decision benchmarks, shipped with its own test suite.

  • OpenDecision - Open alternative: a local semantic decision engine that describes itself as the open-source equivalent of Jev, answering Choice, Noul, and Score questions from structured state and documents without a hosted call.

  • TinyJev - Open alternative: a 596M pointer-head model that answers Choice, Score, and Noul in a single forward pass and returns calibrated confidence meant to be thresholded, so cases it is unsure about escalate to a human instead of being guessed; MLX-first on Apple Silicon, with a System One endpoint and weights on Hugging Face and ModelScope.

  • When a Judgment Layer’s Self-Reported Fields Lie - Independent measurement: tests Jev’s self-reported access-layer fields against ground truth rather than trusting them, reporting a verdict vocabulary reaching three values where the description lists six and a sufficient field that does not separate thin evidence from contradictory evidence; the contradiction reading has no JSON artifact behind it and the write-up says so in its own errata.

  • SemIf - Independent replication: reproduces Jev's typed-decision interface on open models, including an MLX backend on Apple silicon, and measures that typed decisions arrive together while a JSON answer streams token by token.

  • jev-verify - Developer tooling: recomputes Jev's confidence and expected-score identities against outputs published in public repositories rather than live API calls, separating vendor-channel examples (10/10) and recorded responses (843/854) from hand-authored fixtures (115/296), where all 121 outputs whose confidence equals the fractional part of their score are concentrated.

  • AnyJev - Open research: turns open LLMs into Jev-style decision models that read typed decisions and calibrated probabilities from next-token prefill distributions with zero fine-tuning, reducing order-flip rate and calibration error.

  • Verdict - Open alternative: Apache-2.0 118M multilingual bi-encoder that answers Choice, Score, and Noul questions on the same POST /v1/systemone wire format, calibrated with temperature scaling plus a split conformal abstain set with a coverage guarantee (ECE 0.01 to 0.03 on the public suites), runs on CPU or in the browser via ONNX, fits on your own labels in seconds, and its README says it loses to Laya on typed decisions (0.71 vs 0.77).

  • Jev-MedQA - Medical QA: a Jev-style implementation on Qwen3.5-4B that selects answers to text and image questions in one forward pass, reporting 69.42% accuracy versus 67.41% for standard generation with a 10.37x speedup across 153,889 questions from nine medical QA benchmark sets.

  • Jev Prime type: hosted - Text generation: a conversational agent with no language model, where every word is picked from ~4,700 options by Jev Choice questions one at a time, with confidence driving lookahead when the top pick falls below 0.65, beam search across sentence directions, and a self-critique loop that rewrites sentences scoring below threshold; live at talktojev.com, paper at doi.org/10.5281/zenodo.22940945.

  • CLM - Open alternative: an 8B System One model that answers the same Choice and Noul questions behind a TypeSafe-compatible API, matching Jev across computer-use, gaming and tool-calling with up to 9x lower latency and reporting 87.6% on Terminal-Bench 2.1 as a fine-tuned verifier.

  • Bespoke Nimble - Open alternative: a LoRA on Qwen3.5-9B that scores one allowed answer token per Choice, boolean or rubric-score question, released with its data pipeline, training config and eval harness under Apache-2.0, and reporting 90.1% on its 324-example holdout against Jev's 93.2%.

  • NeoHorse-Jev - Open alternative: Apache-2.0 4B decision model from TokenRhythm that answers Choice, Noul and Score questions via prefill-only inference on NeoHorse-1-4B, deployable with vLLM, SGLang or a native Python/CLI/HTTP runtime, scoring 77.70 across six text benchmark groups (highest among open-weight entries with complete results in its published comparison).

  • Jeff - Open alternative: MIT-licensed Qwen3.5 and Gemma 4 fine-tunes answering choice, noul and score on the same /v1/systemone format at about 22 ms per decision, published with a panel that measures Jev itself at 0.828 accuracy and 0.053 ECE while stating it claims no statistical significance.

  • AutoJev - Open recipe: a 27B multimodal decision model trained with full-weight SFT on 73,000 examples over one H200, serving choice, noul and score on /v1/systemone with per-checkpoint provenance and calibration plots released.

  • Lev - Open alternative: a 4B LoRA on a Qwen backbone published as a System One decision model and tagged for calibrated decisions, classification, routing and moderation.

  • openjev - Open implementation: ranks a Doom action menu with one /score call per step and supports per-task fine-tuning, released under MIT with the terminal run recorded.

  • Jev-Omni - Multimodal System One: a Gemma-4-12B fine-tune that answers typed questions over text, images, audio and video, at 1,402 downloads in its first ten days.

  • Bekko System One - Open decision models: independent 17M–400M English models for Choice, Noul, and Score, with public weights, training code, and an ONNX browser demo; v0 remains substantially behind Jev on the project's generalization tests.

  • jevcrypto type: library - Creative experiment: a crypto.randomUUID() look-alike npm package that writes Jev's raw Noul probabilities on 15 code-point permutations of any prompt, plus one Choice for the variant digit, into the bytes of a UUID v4-shaped string; deliberately not cryptographically random.

  • Vev - Open alternative: an open-source Jev implementation with vision input, LoRA fine-tuned on Qwen3.5-4B and 9B, serving Choice, Score and Noul questions on the /v1/systemone wire format with screenshots and photos placed directly in the state so one decision can use both text and image; weights are CC BY-NC 4.0, non-commercial only.

  • WaterSheep - Open alternative: an Apache-2.0 open-weight model that answers noul, choice, score and multi-label questions with a calibrated probability for every option, serves POST /v1/systemone so TypeSafe's Python SDK works after a base-URL change, and has an ONNX build that runs in the browser.

Infra / SDKs / Integrations

Source file: categories/infra-sdks-integrations.md

  • eve - Agent frameworks: Vercel's eve engine ships Jev as the default evaluation model (typesafe-ai/jev) in its experimental evaluate path.

  • AI CLI - Developer tooling: Vercel Labs CLI that can run Jev as the evaluation model for its evaluate command.

  • jev-mcp (jkudish) - MCP ecosystem: proof-of-concept MCP server that puts Jev claim verification, content screening, and candidate ranking behind standard MCP tools.

  • jev-mcp (blakestone-x) - MCP ecosystem: MCP server exposing Jev classify, score, check, match, and screen as tools for any agent, with confidence on every answer.

  • zio-typesafe-ai - Scala ecosystem: ZIO client for TypeSafe AI with a typed DSL over Jev decisions.

  • TypeSafe AI Swift SDK - Swift ecosystem: dependency-free Swift 6 client for Jev Choice, Score, and Noul questions with strict concurrency, configurable authentication and retries, and offline transport tests.

  • laravel-typesafe-jev - PHP ecosystem: unofficial Laravel integration for Jev with typed responses, async requests, scoped dependency injection, and testing fakes.

  • advocaat - Data tooling: small type-safe client for asking Jev questions about a dataset.

  • jevclient - Python ecosystem: async client for Jev published on PyPI.

  • LlamaIndex Jev - Retrieval / RAG: unofficial LlamaIndex adapter where Jev Scores each retrieved passage and Choice/Noul selects the query engine, with nfcorpus nDCG@5 0.340→0.396 at about $0.0003/query.

  • safer-with-jev - Cloud infrastructure: Neon Function proxy for the Neon AI Gateway that routes decisions with Jev.

  • typesafe-ai/skills - Official tooling: installable agent skills package (npx skills add typesafe-ai/skills) that teaches agents the Jev workflow.

  • Smithers - Agent frameworks: TypeScript workflow framework with a Jev session checker wired into its workflows.

  • skillbox - Skills infrastructure: self-hosted versioned skills library that adds optional Jev recommendations using your own TypeSafe or Gateway key.

  • Jevbridge - Agent bridges: ACP and MCP adapter that exposes Jev typed decisions to Codex, Claude, Grok, and other LLMs.

  • jev (Elixir) - Elixir ecosystem: GenServer client that replies with Jev's answer so callers can pattern match on it directly.

  • jev-go - Go ecosystem: community Go SDK for Jev.

  • jev-cli - Developer tooling: small dependency-free CLI for Jev.

  • decide-mcp - MCP ecosystem: configurable decision server with percentage scores and bias-profile routing on top of Jev.

  • typesafe-jev-examples - Starter examples: worked ticket-triage and reranking examples runnable through OpenRouter without an early-access key, shipped with their own sample data and Makefile.

  • ai-python - Python ecosystem: the official Vercel AI SDK for Python carries Jev through its evaluation operation and Gateway examples.

  • Cline plugins - Coding agents: Cline's official plugin collection includes a Jev-driven browser plugin (jev-browser), so Jev arrives as a first-class Cline capability.

  • hono-jev-router - Web frameworks: Hono middleware that routes HTTP requests by meaning rather than by method and path, deciding with Jev.

  • rotom - Local gateways: OpenAI- and Anthropic-compatible API gateway that carries Jev through its model catalog and evaluation path.

  • Jev AI - Developer tooling: public Jev playground and API that puts typed Choice, Score and Yes/No questions to the model about pasted text - ticket triage, moderation, review scoring - and returns a parsed answer with a confidence value in about 0.5 s per decision.

  • jevql - Data tooling: psql-shaped CLI and Go/TypeScript/Python SDKs that run plain SQL on a vanilla Postgres (no extension) and then ask Jev Noul, Choice, or Score questions about each surviving row so the client can apply jev() filters, jev_prob sorts, and jev_choice groups.

  • sqlite-jev - SQLite ecosystem: loadable C extension and Python package that expose Jev Noul, Choice, and Score judgments as SQL functions and batched virtual-table queries with confidence results.

  • jevkit - Developer tooling: Rust CLI that validates Choice/Score/Noul question sets with 13 offline lint rules before any Jev call, then sends the canonical wire payload and prints parsed, confidence-bearing JSON answers to stdout using exit code 2 to reject a billed-but-useless request.

  • jev-use - MCP ecosystem: Claude Code / Codex / pi plugin (MCP server + library, native pi extension) that hands agent steps needing no text output to Jev as typed judgments — untypeable and generation-needing questions are rejected before the call, low-confidence answers come back flagged as priors, and a fail-open PreToolUse gate can only deny or ask.

  • huncho - TypeScript ecosystem: dependency-free SDK that turns Jev Noul, Choice and Score answers into named decisions with enter/exit thresholds (hysteresis), nested decision trees settled in one call, a JSONL journal, replay of a threshold change over recorded answers with no inference, and Brier/reliability calibration, over TypeSafe direct, OpenRouter or Vercel AI Gateway.

  • jev-experiments - Demo collection: 22 latency-focused Jev applications built by Devin, each with its own README and testing notes, spanning shell guards, log sentinels, instant search, reranking, and voice turn-taking.

  • ruby_decision_model - Ruby ecosystem: client for decision models such as Jev, so Ruby applications can put typed questions directly to the model.

  • s1_ruby - Ruby ecosystem: makes System One measurement, and the collapse that follows it, a Ruby primitive, with a TypeSafe provider behind its own spec suite.

  • kojev - Kotlin ecosystem: Kotlin Multiplatform (JVM, Android, iOS) client for Jev that answers Choice and Score questions as the caller's own enums, with one typed way to read answers, no default thresholds, and offline MockEngine tests.

  • hunch - Python and TypeScript ecosystem: libraries that turn Jev Choice, Score, and Noul questions into functions over lists and DataFrames (classify, score, check, where, extract, pick, rank, verify), with request deduplication, caching, and optional escalation of unsure rows to an LLM that must pick from the same labels; TypeScript port at hunch-js.

  • stuntd - Local runtime / learning proxy: Jev-compatible local server on Laya that also proxies a Jev upstream, records every Choice, Score and Noul decision, trains a head per decision site, and answers live with calibrated confidence, falling back to the upstream below its threshold.

  • should-i-jev - Migration tooling: dependency-free CLI that scans LLM logs and code for decision-shaped calls, prices the Jev migration, asks a Jev endpoint which call sites to take (--jev-selfcheck), calibrates typed answers against ground truth (ECE, reliability, risk-coverage), and generates a reviewable migration PR with Choice/Score/Noul map sketches.

  • vellum-assistant - MCP & Integrations: An optional Jev provider in Vellum Assistant sends conversation state and explicit questions to TypeSafe.

  • typesafe-mcp - MCP & Integrations: An MCP server that lets Claude Code, Claude Desktop, Codex and Pi ask Jev typed questions.

  • pi-typesafe - MCP & Integrations: A Pi Jev extension providing a decision tool, terminal playground commands and an API for other extensions.

  • Jevbridge - MCP & Integrations: Exposes a shared structured-decision interface for Jev and other models through ACP, MCP and a CLI.

  • synkora-ai - MCP & Integrations: Synkora includes optional TypeSafe client tools for classification, scoring and yes/no judgments.

  • jev-judge-mcp - MCP & Integrations: Typed judgment tools for MCP agents. TypeSafe's Jev model as verify, screen, find, classify, rerank, decide, compare, extract, review, gate, and score: the model judges, policy decides auto, review, or escalate.

  • plasmallm - MCP & Integrations: A Jev Decisions adapter in a KDE Plasma assistant widget for displaying structured judgments.

  • jevwire - MCP & Integrations: Provides Jev MCP tools, an embeddable library and Claude Code hooks for Agents.

  • harness-router - MCP & Integrations: Fast decision routing for agent harnesses — native MCP with Jev for tool selection and MCTS for multi-step decisions.

  • jev-codex-plugin - MCP & Integrations: Open-source Codex plugin for TypeSafe Jev decision consultation, failure diagnosis, and evidence-based completion review

  • jev-mcp - MCP & Integrations: An MCP server and Claude Code plugin for Jev classification, scoring, checks and batched questions.

  • jev-classifier - MCP & Integrations: A local Jev tool-routing gateway for coding Agents, with an MCP suggestion interface.

  • tenbin - MCP & Integrations: Documentation, an MCP server and a Skill for designing Jev judgments with question linting, batch evaluation and calibration.

  • jev-agent-kit - MCP & Integrations: jevkit: fast typed decisions for agents. CLI and MCP tools (route, triage, guard, grep, rank, compact, judge) on TypeSafe Jev. Zero dependencies.

  • Jev-AI-Skill - MCP & Integrations: One AI skill + MCP server for Claude Code, Codex and Hermes: Jev (TypeSafe) gates large-model turns (event triage, review verdicts, owner questions, tool choice), runs build-and-repair loops with independent review, routes models and skills, and guards Git steps. Script-only watch mode and a shared HTTP server.

  • jev-as-quant - MCP & Integrations: Typed System-1 decisions (Laya/Jev) as the judgment layer of a quant research stack, with Claude as System 2. Requirements → design → code → experiments.

  • jev-mcp - MCP & Integrations: An evaluation-focused Jev MCP server for individual questions, batch processing and question or threshold comparisons.

  • jev-mcp-server - MCP & Integrations: MCP server for Jev (TypeSafe System One): the three official question types — choice, score, noul — plus batch classify. Calibrated probabilities, ~0.5s, <$0.001/call.

  • jev-skill-router - MCP & Integrations: Keep skill catalogs outside the main LLM context. Jev selects relevant skills through one read-only MCP tool.

  • jev-workbench - MCP & Integrations: Defines, tests, and publishes Jev decision functions in a local UI so backends and Agents can call fixed versions.

  • jev-agent-toolkit - MCP & Integrations: Jev-first portable Agent Skill and optional MCP bridge for Claude Code, Codex, Cursor and compatible agents.

  • jev-mcp - MCP & Integrations: MCP server wrapping TypeSafe's Jev System One models — typed noul/choice/score judgments for AI agents

  • jev-mcp - MCP & Integrations: Local MCP server exposing TypeSafe Jev (System One decision model) as native Claude Code / Codex tools

  • jev-mcp - MCP & Integrations: MCP server for Jev (TypeSafe AI's System One model) — give any agent typed, calibrated decisions: classify, score, check, gate risky tool calls. Try free: jevtypesafeai.com

  • jev-mcp-spring - MCP & Integrations: On success each tool returns its typed result directly. On failure the tool call fails at the MCP protocol level (`isError: true`) with a short, safe message — no response body or stack trace is ever echoed back.

  • jev-playwright-mcp - MCP & Integrations: Jev-augmented Playwright MCP proxy — page-state triage, prompt-injection shielding, goal-based snapshot pruning, risky-action gating. Drop-in wrapper around @playwright/mcp for any coding agent.

  • jev-rust-review - MCP & Integrations: Rust-aware code review for Claude Code and coding agents, powered by TypeSafe Jev

  • jev-toolkit - MCP & Integrations: MCP-first toolkit for TypeSafe/Jev — the System One decision model. One stdio server (jev mcp) serves any MCP-capable harness, backed by one local event log and Prometheus impact metrics you can scrape into your own Grafana.

  • n8n-nodes-jev - MCP & Integrations: Jev by TypeSafe AI for n8n: typed decisions, probabilities, and confidence-aware workflows

  • n8n-nodes-typesafe-jev - MCP & Integrations: An n8n community node for submitting typed questions to TypeSafe Jev.

  • toolJev - MCP & Integrations: Code Mode for MCP, where the sub-model is a calibrated decision model (Jev), not an LLM. Benchmarked on MCPToolBench++, LiveMCPBench, When2Call and live Claude agents.

  • typesafe-jev-mcp - MCP & Integrations: This repository is an MCP server for TypeSafe Jev that provides a single evaluate tool taking state plus typed questions and returning noul, choice, or score answers with probabilities.

  • typesafe-jev-opencode - MCP & Integrations: Jev is not a conversational replacement for Gemini, Claude, or GPT. It evaluates application state against typed questions and returns structured answers and probabilities that an agent can use to route or gate work.

  • jev_ampcode - MCP & Integrations: An Amp plugin for comparing supplied alternatives against supplied evidence and priorities.

  • jev-eyes - MCP & Integrations: Also in the state: `image` (size, source), `blocks` (`[x, y, w, h]` boxes with OCR confidence) and, if installed, `labels`. `see(img, compact=True)` keeps only `image`, `text` and top label names when tokens matter more than positions.

  • jev-in-mcp - MCP & Integrations: MCP relay that adds use_jev to every server: Jev picks the tool calls, the calling model writes the values Jev cannot choose, the relay executes. Built on jev-dev-kit.

  • jev-mcp - MCP & Integrations: MCP local que expone Jev (TypeSafe) como herramienta para Claude Code, Codex, Hermes y cualquier agente: ask_jev y list_jev_models, sin dependencias

  • jev-mcp - MCP & Integrations: A Rust MCP server for TypeSafe AI Jev structured decisions

  • jev-mcp - MCP & Integrations: `jev-mcp` exposes TypeSafe AI's Jev decision model as four conservative, read-only MCP tools for **bounded probabilistic decisions**.

  • jev-mcp-open-source - MCP & Integrations: Self-hosted Jev MCP on Cloudflare Workers with intent routing, retrieval reranking and batch judgments

  • jev-review-mcp - MCP & Integrations: Single-purpose MCP server (one tool, one job): a code-review gate powered by TypeSafe Jev (System One decision model).

  • jev-routing - MCP & Integrations: Go Jev harness for Claude Code, Codex, and Grok Build. No npx. Not an MCP server.

  • jev-screen-mcp - MCP & Integrations: Single-purpose MCP server (one tool, one job): a content-moderation gate powered by TypeSafe Jev (System One decision model).

  • jev-tool-search - MCP & Integrations: Tool search for LLM agents: BM25 vs embeddings vs rerankers vs Jev on 525 real MCP tools, plus an experimental Jev search engine

  • jevmod - MCP & Integrations: Moderation for communities and apps: every message gets a probability for **spam, scam, harassment, nsfw, off-topic, self-harm, doxxing, sexual content involving minors**, and for **rules you write in plain English**. You set the thresholds and the actions. Every decision is logged with its numbers.

  • mcp-server-jev - MCP & Integrations: Typed AI decisions for Codex, Claude and any MCP client, powered by TypeSafe Jev. Classify, score and evaluate with one generic tool.

  • openclaw-typesafe-ai - MCP & Integrations: An independent OpenClaw plugin registering one explicitly invoked typesafe_decide tool.

  • openrouter-jev-mcp - MCP & Integrations: A Python decision gateway and Model Context Protocol (MCP) server exposing TypeSafe's Jev model through OpenRouter's decisions endpoint to Claude Code, Codex, and Cursor agents.

  • composio - SDK & Decision Frameworks: An optional TypeSafe provider for Composio that uses Jev to choose among tools and bounded argument options.

  • ai - SDK & Decision Frameworks: The TypeSafe provider in AI SDK lets TypeScript applications call Jev through the shared evaluate interface.

  • eliza - SDK & Decision Frameworks: An optional TypeSafe HTTP adapter in Eliza’s source, not registered with the Agent runtime by default.

  • langchainjs - SDK & Decision Frameworks: An optional LangChain.js TypeSafeClassifier integration for sending state and predefined questions to Jev.

  • rig-typesafeai - SDK & Decision Frameworks: An experimental TypeSafe crate in Rig for expressing Jev questions and answers with Rust types.

  • jev - SDK & Decision Frameworks: jevos is an open-source, Jev-compatible alternative to TypeSafe's Jev for yes/no decisions that runs entirely on a laptop CPU. Send text plus a yes/no question over HTTP and get back P(yes) from a single forward pass of a 1B model — no text generation, no GPU required.

  • req_llm - SDK & Decision Frameworks: A TypeSafe provider for calling Jev through ReqLLM’s evaluate interface in Elixir.

  • simple-jev - SDK & Decision Frameworks: Adapter turning open LLM endpoints into Jev-compatible classification services without training a separate classifier head.

  • jev-skill - SDK & Decision Frameworks: This project is a collection of Jev use cases, workflows, and agent skills with a stdlib-only Python decision wrapper.

  • openjev - SDK & Decision Frameworks: An independent System One decision server compatible with Jev’s API, running an open DiffusionGemma model.

  • instructor-php - SDK & Decision Frameworks: A TypeSafe Decision driver within Instructor PHP’s Polyglot module.

  • jev-visual - SDK & Decision Frameworks: Local Jev-like visual inference experiment on Apple Silicon Mac. Scores and classifies single images across multiple questions with 3 playable game demos.

  • jev-dsh-decision - SDK & Decision Frameworks: Provides Jev structured decision support for Agent Harness to recommend tools, Skills and Agents and return judgments with probabilities, with a native DeepSeek Harness plugin and an iPolloWork entry serving OpenCode, DeepSeek Harness and Codex Harness.

  • pi-fabric - SDK & Decision Frameworks: Pi’s programmable runtime includes an optional Jev loop for observing state, making decisions and running bounded actions.

  • typesafe-sdk-js - SDK & Decision Frameworks: The JavaScript and TypeScript SDK published by TypeSafe, with typed Jev requests and answers.

  • openai-scala-client - SDK & Decision Frameworks: A dedicated TypeSafe module in a Scala client that supports multiple AI providers.

  • typesafe-sdk-python - SDK & Decision Frameworks: Official TypeSafe Python SDK with synchronous and asynchronous clients for Jev System One, plus question and answer types.

  • jevbench - SDK & Decision Frameworks: JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.

  • runline - SDK & Decision Frameworks: A TypeSafe plugin exposing Jev decisions as callable actions in Runline Agent JavaScript.

  • ai - SDK & Decision Frameworks: A Jev forwarding endpoint in the Hack Club AI proxy, using its authentication, limits and usage logging.

  • effect-agent - SDK & Decision Frameworks: An Effect Agent TypeSafe decision provider for typed question sets and optional model selection.

  • jeview - SDK & Decision Frameworks: An unofficial local visualizer for Jev (TypeSafe): a live view of every call your code makes. Not affiliated with TypeSafe AI.

  • Jev - SDK & Decision Frameworks: Unofficial TypeSafe Jev showcase — System One decisions, not chat.

  • solar-mini4-jev - SDK & Decision Frameworks: A drop-in wrapper that exposes Upstage **Solar Mini4** through the TypeSafe Jev System One API shape.

  • ask-jev-skill - SDK & Decision Frameworks: A skill for Hermes and other agents to query TypeSafe Jev for bounded option judgments and confidence escalations.

  • go-jev - SDK & Decision Frameworks: Go SDK and CLI for TypeSafe Jev: typed decisions (yes/no, choice, score) from a model

  • jev-capability-atlas - SDK & Decision Frameworks: This repository collects real Jev API-call receipts, test suites, and bilingual guides to map which narrow-decision tasks suit Jev and how Agents should evaluate and report fit.

  • minojev - SDK & Decision Frameworks: Decisions, not tokens: minojev reads calibrated, typed probability distributions straight from hidden states in one forward pass — zero output tokens, fully reproducible on a laptop CPU.

  • jev-agent-design-with-topk-logits-choices - SDK & Decision Frameworks: Research design for a Jev-native agent system: tool integration, speculative parameter proposals, external helper logits Top-k proposals with Jev-controlled fallback ,decision-aware hierarchical memory, and dependency-aware replanning.Feature:Jev naturallanguage conversation prototype using external helper logits and dynamic Top-k token selection.

  • jevalyn - SDK & Decision Frameworks: The decision layer for your Rails app. A Rails-native wrapper around TypeSafe's Jev System One API: typed, calibrated decisions in your control flow.

  • jev-to-answer - SDK & Decision Frameworks: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • swift-jev - SDK & Decision Frameworks: A Swift client for TypeSafe AI's Jev — typed judgements, not text

  • swift-typesafe - SDK & Decision Frameworks: A community Swift TypeSafe client with typed questions, dynamic questions and response parsing.

  • discern - SDK & Decision Frameworks: A semantic control-flow library for Effect: Jev's `Choice` / `Noul` / `Score` answers become typed branches. Thresholds are caller-supplied, and an answer below them takes an explicit `Uncertain` branch the compiler forces you to handle rather than being rounded up to the top label. Procedure routing filters candidates with deterministic predicates first and skips the model call entirely when one candidate survives.

  • jev-rs - SDK & Decision Frameworks: System One judgments (noul/choice/score) from any LLM in one prefill — a Rust, Jev-compatible /v1/systemone engine

  • typesafe-ai - SDK & Decision Frameworks: A Rust TypeSafe client with asynchronous reqwest or blocking ureq backends and observable retries.

  • learn-jev-end-to-end - SDK & Decision Frameworks: **Learn Jev end to end** is a free, hands-on course. In 12 short notebooks you go from *"what is Jev?"* to building **13 real AI tools** with it: an email triage job, a scam-text detector, a code vulnerability hunter, an agent safety guard and more. You need **one API key**, and running the whole course costs **less than $0.20**.

  • typesafeai-dotnet-sdk - SDK & Decision Frameworks: Community .NET SDK for TypeSafe AI and Jev, providing asynchronous typed evaluation for Choice, Score, and Noul primitives.

  • jev-dspy-lab - SDK & Decision Frameworks: Reproducible calibration and selective-risk benchmarks for Jev/TypeSafe decisions in DSPy workflows

  • typesafe-sdk-java - SDK & Decision Frameworks: A community Java TypeSafe client with a Spring Boot Starter for configuring Jev calls.

  • jev-ai-sdk-form-router - SDK & Decision Frameworks: Route form submissions to the right people with Jev and AI SDK.

  • SpecPi - SDK & Decision Frameworks: A Pi configuration and extension bundle with an optional Jev advisor for capabilities and workflow checks.

  • typesafe-sdk-go - SDK & Decision Frameworks: A Go TypeSafe SDK for defining typed questions and reading Jev choices, scores, and probabilities.

  • zod-jev - SDK & Decision Frameworks: Adds semantic rules to Zod validation, such as checking whether text matches a description or contains personal information.

  • hermes-jev-plugin - SDK & Decision Frameworks: TypeSafe Jev (System One) decision tools for Hermes Agent: jev_check / jev_route / jev_score / jev_evaluate

  • jev-dsl - SDK & Decision Frameworks: An early Haskell DSL that describes labeled Jev questions, renders requests and decodes matching answers.

  • jevgo - SDK & Decision Frameworks: Unofficial Go client for the TypeSafe AI System One API (Jev) — typed questions in, calibrated answers out.

  • JevOps - SDK & Decision Frameworks: Jev is a **gate**, not a generator. This package does **not** write Lean. Lake (or another oracle) lives in the implementation that *uses* the kernel.

  • typesafe-sdk - SDK & Decision Frameworks: A community Ruby client for TypeSafe System One, defaulting to jev-latest.

  • daf-jev - SDK & Decision Frameworks: A Python toolkit for Jev requests, batch evaluation, calibration and MCP access.

  • jev_jsonschema - SDK & Decision Frameworks: `probabilities` is keyed by your schema's values, not Jev's internal labels, so a score of `1`–`5` reads as `"1"`–`"5"` and not `"0"`–`"4"`. Noul questions carry no confidence of their own, so `confidence` is `None` for booleans and numbers.

  • open-bonsai-jev - SDK & Decision Frameworks: openjev's mechanism, Bonsai's weights: typed decisions read straight from one forward pass of a 1.75-bit 27B model. Credit to TheoLeeCJ (SemIf/OpenJev) and PrismML.

  • typesafe-ai-rs - SDK & Decision Frameworks: An independently maintained Rust SDK with async and blocking clients, retries and response metadata.

  • jev-android - SDK & Decision Frameworks: A Kotlin Android SDK for UI automation powered by TypeSafe Jev, with an accessibility runtime and sample app.

  • jev-go - SDK & Decision Frameworks: A Go TypeSafe System One client with typed questions, answers and batching helpers.

  • jev-sdk-java - SDK & Decision Frameworks: Type-safe Java 21 client for the TypeSafe AI Jev (System One) decision API

  • jevriel - SDK & Decision Frameworks: **Codex · Claude Code · Portable skill** | [Apache-2.0 code and docs](LICENSE) | Early release

  • typesafe_sdk - SDK & Decision Frameworks: An Elixir TypeSafe SDK that brings Jev’s typed questions and probabilistic answers into Elixir applications.

  • usejev - SDK & Decision Frameworks: Run Laya locally with Bun: native ONNX inference, a TypeSafe-compatible API, and a bilingual decision playground.

  • goodall - SDK & Decision Frameworks: An optional TypeSafe package in a Go Agent library, using Jev as a tool or routing judge alongside chat models.

  • jev - SDK & Decision Frameworks: A Go client for the TypeSafe AI's System One API and its model, Jev.

  • jev-java - SDK & Decision Frameworks: Unofficial Java SDK for TypeSafe Jev and Vercel AI Gateway, with Spring Boot and WebClient support

  • jev-pilot - SDK & Decision Frameworks: Fast System-1 Decision, Arbitration & Safety Engine for Autonomous AI Agents (Powered by TypeSafe Jev)

  • jevlang - SDK & Decision Frameworks: The simplest way to write decision workflows in Python. Python with a smart if.

  • questions - SDK & Decision Frameworks: A TypeScript decision library that asks typed questions via Zod or native batches, defaulting to TypeSafe Jev, with optional Vercel or generative adapters.

  • ask-jev-ai - SDK & Decision Frameworks: Most AI demos generate text. Jev does not. It reads a sentence and returns typed answers with probabilities: a choice, a yes or no, a score. That makes it usable as a primitive inside ordinary code rather than a chatbot bolted onto a page.

  • jev-does-not-play-dice - SDK & Decision Frameworks: Code, recorded outputs, and analysis scripts for probability-output experiments with Jev: fair random draws, Noul (Yes/No) questions, and forecast documents.

  • jev-layer - SDK & Decision Frameworks: Portable System-1 decision layer for agent harnesses with host-owned routing, receipts, replay, and fail-open integrations.

  • jev-numeric - SDK & Decision Frameworks: **Both are multiway decision trees; decimal-digit decoding is a ten-way instance.** On an aligned decimal grid, they can have identical branches and leaves, expressed through different prompts. The digit is a **Choice option**, not a vocabulary token; Jev returns the option probabilities directly.

  • Jev4Mellea - SDK & Decision Frameworks: This is an unofficial, synchronous adapter. Jev evaluates text; it does not generate or repair it.

  • jevex - SDK & Decision Frameworks: An Agent experiment where Jev directs a tool loop, a chat model fills arguments and prose, and MCP tools execute.

  • typesafe-ai-rails - SDK & Decision Frameworks: Ruby on Rails integration gem for TypeSafe AI and Jev, providing model-level classification and decision policy patterns.

  • typesafe-sdk-rust - SDK & Decision Frameworks: A Rust client for TypeSafe with asynchronous and optional blocking calls plus typed question and answer wrappers.

  • dsh-jev-decide - SDK & Decision Frameworks: This DSH plugin registers TypeSafe Jev as an Agent tool that returns calibrated probabilities for noul, choice, and score judgments without generating text.

  • everything-about-jev - SDK & Decision Frameworks: tell you everything about jev,TypeSafe AI's System One model for typed decisions.

  • jev - SDK & Decision Frameworks: Ruby client for the typesafe AI Jev model

  • jev-benchmark - SDK & Decision Frameworks: Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs Claude Sonnet 5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.

  • jev-builder - SDK & Decision Frameworks: A browser form for building requests to TypeSafe's Jev: pick a template, fill in the blanks, copy the request. No JSON, no install, runs locally.

  • Jev-Persian-Benchmark - SDK & Decision Frameworks: Benchmarks for **Jev** on **480 authored general Persian questions** and a **24-excerpt classical Persian poetry pilot** (48 main questions plus 48 controls). Related general questions are batched; poetry questions run individually. Raw responses are saved and answers are scored locally, without a runtime model judge. The general benchmark's historical Laya comparison is retained below.

  • Jev-PhoneControl - SDK & Decision Frameworks: Visual Android automation powered by three agents: vision, a text-only supervisor, and TypeSafe JEV. Executes actions through ADB with a local web console.

  • jev-skills - SDK & Decision Frameworks: Practical agent skills and examples for building with Jev. API setup, routing, ranking, and evidence checks.

  • jev-web-analyzer - SDK & Decision Frameworks: See what Jev thinks about your SaaS website — powered by ReplyNodes web context and Vercel AI Gateway.

  • jevclient - SDK & Decision Frameworks: An asynchronous Python Jev client that batches typed questions in one request.

  • jevgo - SDK & Decision Frameworks: A community Go client with a standard-library core and optional Langfuse tracing.

  • jevrag - SDK & Decision Frameworks: Replaces hardcoded RAG thresholds with explicit calibrated decision points. Five primitives (retrieval stopping, chunk splitting, context selection, answer abstention, cache trust) behind one swappable state → Decision → confidence → action interface, each evaluated on real datasets with a calibration harness that reports honestly.

  • qualm - SDK & Decision Frameworks: A TypeScript wrapper for Jev decisions with an explicit unsure branch.

  • typesafe-go - SDK & Decision Frameworks: An unofficial Go client without third-party dependencies for System One calls and model discovery.

  • typesafe-sdk-php - SDK & Decision Frameworks: A community TypeSafe SDK for PHP 8.2+, with synchronous calls and Guzzle-based async requests.

  • jev - SDK & Decision Frameworks: Unofficial Go client for TypeSafe's System One API and its model, Jev.

  • jev_dart - SDK & Decision Frameworks: Jev Dart SDK to build cli, server and Flutter apps.

  • jev-architecture-research - SDK & Decision Frameworks: Black-box reverse engineering research archive for the Jev decision model

  • jev-by-example - SDK & Decision Frameworks: Ten runnable Jev examples for agent decisions: memory conflicts, tool-result checks, recovery, context selection, and handoffs. JavaScript, zero dependencies.

  • jev-doom - SDK & Decision Frameworks: Watch Jev play Freedoom in a local dashboard. TypeSafe direct and Vercel AI Gateway, inspectable decisions, and bounded spending.

  • jev-go - SDK & Decision Frameworks: A small unofficial Go SDK for Jev calls and model listing.

  • jev-phone - SDK & Decision Frameworks: Drive a phone with a model that never writes a word. TypeSafe's Jev picks each action, phone-use runs it on iOS and Android.

  • jev-starter - SDK & Decision Frameworks: TypeScript patterns for thresholds, fallbacks, human review and evaluation on top of the TypeSafe SDK.

  • typesafe-ai-jev-example - SDK & Decision Frameworks: This repository provides six runnable Python demos and four notes covering TypeSafe Jev primitives and composition patterns, with offline mock mode and committed live samples from jev-1.13.0.

  • typesafe-go - SDK & Decision Frameworks: A TypeSafe System One client that uses only the Go standard library to send Jev questions and read structured answers.

  • typesafe-go - SDK & Decision Frameworks: An unofficial Go SDK with typed answers, retries and context cancellation.

  • typesafe-rs - SDK & Decision Frameworks: A community Rust Jev client with async requests, an optional blocking interface and local mock testing.

  • claude-jev-mod - SDK & Decision Frameworks: Typed decisions in Claude Code: adds $.jev over TypeSafe's Jev, through OpenRouter, Vercel AI Gateway, Cloudflare Workers AI, LiteLLM or the TypeSafe API.

  • jev_playground - SDK & Decision Frameworks: Jev answers typed questions about a state with calibrated numbers. This repo is where we find out which of those numbers deserve to drive code, and where a regex or a constant does the job better.

  • jev-go-sdk - SDK & Decision Frameworks: Dependency-free Go client for TypeSafe AI's System One API and the Jev model

  • jev-is-not-odd - SDK & Decision Frameworks: A probabilistic, AI-powered utility to determine if a number is not odd (or not even) using TypeSafe's Jev model and the Vercel AI SDK.

  • jev-lab - SDK & Decision Frameworks: This repository provides a single-page classifier demo that sends text with choice questions to Jev and displays the request JSON, probability distribution, confidence, latency, and token usage.

  • jev-lab - SDK & Decision Frameworks: Hands-on research lab for TypeSafe's Jev (System One model): reproducible benchmarks of Noul/Choice/Score primitives, confidence gating, fan-out latency, agent control — plus a living audit of the Jev ecosystem.

  • jev-labs - SDK & Decision Frameworks: Never confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong verdicts. Film, code, and every captured call.

  • jev-msw - SDK & Decision Frameworks: Mock Jev API decisions with MSW for deterministic tests without real API calls or credits.

  • jev-swap - SDK & Decision Frameworks: Find the LLM calls in your codebase that are really decisions, see what they'd save on TypeSafe Jev, and prove it on live traffic before you swap. TypeScript, JavaScript, Python.

  • jev-tab-order - SDK & Decision Frameworks: **Organize the entire window with a single Jev API request.** Grouping and ordering decisions are evaluated together, regardless of the number of tabs.

  • jev-triage - SDK & Decision Frameworks: Automated issue & PR triage for open-source maintainers, powered by Jev (TypeSafe AI).

  • jevcore - SDK & Decision Frameworks: A judgement primitive for TypeSafe **System One / Jev** — ask N things × K typed questions in bounded, cheap, fail-open requests — plus the **JEV harness**, the closed boundary in code that makes a Jev answer safe to consume. Standard library only.

  • jevinf - SDK & Decision Frameworks: An inference engine for decision models of the Jev kind: each candidate path runs as segmented forwards with prefix reuse, and the Jev wire contract is served on top. NanoJev is the backend wired up today.

  • jevish - SDK & Decision Frameworks: Every mode auto-curries when called with only the patterns:

  • jevpolicy - SDK & Decision Frameworks: JevPolicy is an open-source TypeScript decision runtime that turns probabilistic judgments from Jev, accessed through Vercel AI Gateway, into versioned, deterministic, replayable, observable application decisions.

  • Jevs-Garage - SDK & Decision Frameworks: System One turns unstructured or structured state into fast probabilistic judgments. Instead of asking for free-form prose, these demos ask `Choice`, `Score`, and `Noul` questions and receive typed values with uncertainty that application code can reason about.

  • typesafe-ai-ruby - SDK & Decision Frameworks: A stdlib-only Ruby client that sends Choice, Score, or Noul questions to TypeSafe System One.

  • typesafe-sdk-swift - SDK & Decision Frameworks: An experimental Swift SDK for TypeSafe using Swift Package Manager, Swift concurrency, and URLSession.

  • TypeSafeSDK - SDK & Decision Frameworks: An unofficial .NET client that POSTs state and typed questions to TypeSafe /v1/systemone. The parent repo also contains unrelated SDK dumps.

  • langchain - SDK Integrations: An optional Jev classifier integration for Python LangChain workflows.

  • pydantic-ai - SDK Integrations: An optional TypeSafe provider and Jev model integration for Pydantic AI.

  • ax - SDK Integrations: Ax provides a TypeSafe integration for boolean or finite-class signatures and native Jev answers.

  • ruby_llm-typesafe - SDK Integrations: A TypeSafe provider for RubyLLM 2 that exposes Jev’s three judgment types through structured output.

  • jev-resilience - SDK Integrations: A Spring WebFlux integration that detects error messages hidden in HTTP 200 response bodies.

  • laya-mlx - Local runtime: independent MLX port of the Laya checkpoints that runs typed decisions natively on Apple Silicon — 13.4 ms median end-to-end per short English decision, 7.4 ms with the multilingual checkpoint, and zero output tokens, with no PyTorch, Transformers runtime, or cloud API.

  • laya-Ascend - Local runtime: Ascend NPU fork of the Laya checkpoints that answers the same Choice, Score and Noul questions on Huawei 910B hardware — 37–47 ms median for a four-question request, 33.8x–70.9x faster than the same request on a single container CPU thread, with an output-equivalent SDPA decision head that avoids torch_npu's CPU fallback on aten::_transformer_encoder_layer_fwd.

  • jev-trust - Python ecosystem: trust middleware for the Jev API that logs every typed decision, measures calibration in your own domain from outcomes you record (accuracy, Brier, top-label ECE, C = 1 − ECE), annotates each answer with its measured effective confidence, fires overconfidence alerts, and signs the evidence (ed25519) for independent recomputation.

  • JarvisCore - Agent frameworks: Python multi-agent runtime that ships Jev natively from 1.12, where agents ask typed Choice, Score and Noul questions through a decision client separate from the text model, the Kernel picks a specialist subagent by Choice, and each retrieved RAG passage is withheld from the generating model when its prompt-injection Noul exceeds 0.70.

  • hunch (carldaws) - Ruby ecosystem: turns judgment calls into control flow — if Hunch.likely?("fraudulent", given: order) reads like plain Ruby but branches on a typed Jev answer, with pick for Choice, rate for Score, and graded predicates from possibly? to almost_certainly?; a TypeScript port offers the same interface.

  • Early experimentation using Jev to rethink harness UX - Harness integration: an agent platform wires Jev into its LLM harness as a callable tool for search, approvals and context, reporting 2,000 expense reports categorized in 20 seconds for five cents.

  • jev-mcp (burnigtm) agent: Multi - MCP ecosystem: server that puts Jev into the coding loop for Cursor, Codex, and any MCP client, with 20 test files behind it.

  • jev-architect - Design skill: finds, designs, and evaluates Jev decision loops, packaged as a skill with references on decision design and delivery.

  • Building a Harness with Jev - Framework guide: LangChain's walkthrough of wiring Jev into an agent harness as the decision layer, from a team that then published its own evaluation of Jev as a judge.

  • system-one-adapter-python - Python ecosystem: TypeSafe AI's official open-source drop-in adapter for running and benchmarking Jev System One decision evaluations across OpenAI- and Anthropic-compatible LLM APIs.

  • neurolink - Provider abstraction: the pipe layer of an AI nervous system — Juspay's TypeScript interface connecting provider neurons to an application, with decide as a first-class inference type alongside generate and stream.

  • jev-spring-boot-starter - Java ecosystem: Spring Boot 4 starter that puts Jev behind Spring MVC and RestClient.

  • mysql-ailike - Database filtering: MySQL plugin that filters rows by a natural-language predicate instead of a literal one, powered by Jev.

  • FastJev - Local runtime: self-hosted Python SDK and System One-compatible API for runtime-defined Choice, Boolean, and Score decisions on pinned open models across Torch, vLLM, MLX, llama.cpp, and WebGPU, with committed row-level benchmarks and checksums.

  • Search with Jev and Milvus - Search engineering: nine runnable notebooks combine Gemini embeddings and Milvus retrieval with Jev Noul and Choice judgments, while Python applies ranking, filtering, routing, and stopping policies to synthetic examples.

  • jeff (logan-markewich) - Self-hosted runtimes: self-hosted drop-in replacement for TypeSafe Jev powered by GliFormer, exposing native Choice, Score, and Noul decision endpoints without cloud API dependencies.

  • CloJev - Clojure ecosystem: unofficial portable Clojure SDK for System One, so Clojure applications can put typed questions to Jev without a Java interop layer.

  • spring-ai-typesafe type: library - Java ecosystem: Java SDK for TypeSafe AI's Jev API and Spring AI integration, providing typed decisions for evaluation as a judge, guardrails, and RAG post-processing.

  • JevFlow type: library - TypeScript ecosystem: composes Jev Noul, Score, and Choice decisions into deterministic threshold workflows that batch into a single systemOne call and return an ordered, explainable action set instead of side effects, with matched rules recording the actual value behind each action and a mock provider so policy tests run without an API key.

  • Jev AI Tools type: hosted - Developer education: hosts six bounded recipes plus a custom builder for AI SDK Choice, Score, and Boolean evaluations; recipes display answer probabilities separately from provider confidence and use deterministic local thresholds to pause uncertain routes for review, while every configuration exports as TypeScript.

  • jev-foundation-models - Apple platforms: a Swift 6 bridge that runs Jev decisions through Apple's Foundation Models on device, with a protocol-based model interface and six tests.

  • cu-Jev - Inference engine: a CUDA-native implementation of the Jev System One API that keeps decisions GPU-resident, shipping a Starfighter demo and a benchmark script.

  • jevcache - Cost control: memoizes Jev-class decisions so a repeated question is served from cache instead of a new call, keeping repeats deterministic and free.

  • Qwev type: self-hosted - Local inference: turns dense Qwen3 and Qwen3.5 checkpoints into a training-free Jev-style Noul, Choice, and Score service that shares one state prefill across isolated questions and, on its included 27-question Qwen3.5-9B/A100 fixture, reports 0.500 s versus 13.554 s for generated JSON.

  • jev-sdk-go type: library - Go ecosystem: dependency-free Go 1.24+ client for Jev Noul, Choice and Score questions that reads Choice and Score answers back as the caller's own types, rejecting any label or level the question never offered, with retries, OpenRouter support, and eleven examples tested against an in-process fake of the API.

  • jevcompat type: cli - Interoperability: a 48-requirement spec of the POST /v1/systemone wire contract, each rule citing TypeSafe's docs, OpenAPI file or SDKs, and a suite that checks any Jev-compatible server against it (Choice probabilities keyed by option and summing to 1, Score equal to Σ i·p, 2–255 options, error shapes, answers that stay put when question ids or order change), finding 2 of the 8 most-starred open ports conformant, with a reference mock that breaks each rule on purpose and a proxy that fixes what can be fixed.

  • typesafe-ai-php - PHP ecosystem: A modern TypeSafe AI client and SDK, with result classes and classic requests

  • Jev type: library - PowerShell ecosystem: PowerShell module for building Jev Noul, Choice, and Score questions and returning named answers as pipeline-friendly properties.

  • djev-run - Serving: deploys DiffusionGemma-Jev behind a TypeSafe-compatible API on a Cloud Run GPU with snake, dino and tetris demos wired to the decision endpoint.

  • jev-symfony-bundle type: library - PHP / Symfony: Symfony bundle providing typed Jev clients, validation constraints (#[JevNoul], #[JevChoice]), Workflow guards, and WebProfiler panels.

  • JevT++ type: library - C++ integration: independent C++20 library with compile-time enum schemas, typed Choice/Noul/Score results and abstention, local Laya inference through ONNX Runtime or ggml, and an opt-in TypeSafe System One HTTP backend tested with mocks and loopback HTTP rather than live-provider calls.

  • Sim - Agent frameworks: open-source collaborative workspace for building, deploying, and monitoring AI agents featuring native TypeSafe System One evaluation and decision provider integration.

  • RubyLLM type: library - Ruby ecosystem: official Ruby gem connecting TypeSafe judgment models to RubyLLM with a native System One protocol for typed questions, probabilistic answers, and error normalization.

  • jev-style agent: Multi type: self-hosted - Local runtime: pip install "jev-style[torch]" (or [mlx] on Apple silicon) serves an open 0.8B Qwen3.5 decision model behind a System One-compatible /v1/systemone API that answers Choice, Score, and Noul questions with calibrated probabilities in one pass over inputs up to 25,600 tokens (0.15–0.2 s per short request with MLX on an M1 Max, after the first call), and ships a Claude Code guard that turns four Noul checks and a risk Score into allow, ask, or deny through code-owned thresholds.

  • grev type: cli - Developer tooling: grep, sort, cut and uniq that match by meaning — grev 'is a vegan meal' menu.txt keeps the lines whose Jev Noul clears 0.5, sibling filters route by Choice and rank by Score, and the output is always your own input, never generated text.

  • decision-gate type: library - Cost and rate control: npm library that every Jev request in a loop goes through, which waits for room under 80% of the account's requests-per-minute and tokens-per-second limits, pauses every caller sharing the account, across processes, for the server's Retry-After delay when the service answers 429, refuses any request that would pass a per-key daily spend ceiling (USD 0.20 by default), and keeps an opt-in cache of answer probabilities under caller-chosen keys, so a repeated question over the same state is not paid for twice.

  • ollaya type: self-hosted - Local runtime: serves open decision models behind a wire-identical /v1/systemone endpoint, so an existing Jev client only has to point TYPESAFE_BASE_URL at the daemon, and reports its recommended model at 0.722 accuracy against Jev's 0.738 on typed decisions.

  • Tiltmeter type: proxy - Monitoring: drop-in /v1/systemone proxy and Pydantic AI client that records every Jev answer's probabilities and alerts, without labels, when jev-latest switches versions, a question's answers drift (chi-square-tested PSI), answers crowd a decision threshold, or estimated accuracy falls.

  • DecisionKit type: library - .NET ecosystem: provider-independent .NET decision engine whose domain package holds no Jev URL, header or DTO, mapping Choice, Score and Noul questions onto POST /v1/systemone from a separate provider package, with a runnable ASP.NET ticket-triage sample that picks the owning team and escalates to a human at a normalized Score of 0.8, and 1,096 tests across net8.0 and net10.0 that run with no HTTP.

  • intern-decision-mlx type: self-hosted - Local runtime: serves Shanghai AI Lab's open Intern-Decision-0.8B vision decision model on Apple Silicon behind a /v1/systemone-shaped endpoint that also takes screenshots, so an agent can threshold the calibrated Choice, Score and Noul answers and escalate the rest — 0.9 s per 1080p screenshot downscaled to 1024 px on an 8 GB M2 MacBook Air, and the same answers as the lab's PyTorch reference on 67 of 67 fields.

Game & Simulation

Source file: categories/game-simulation.md

  • typesafe-mario - Gaming: TypeSafe/Jev agent that plays Super Mario Bros. from structured emulator state, choosing each action from emulator-derived features.

  • jev-drone - Robotics simulation: camera-only autonomous drone in MuJoCo that puts a Jev judgment model in the control loop at 2.5 Hz.

  • tsai-sc - Gaming: drives original StarCraft shareware through keyboard and mouse with Jev action probabilities recorded per decision.

  • jev-plays-pokemon - Gaming: reads Pokémon Red game state as text, answers typed questions each turn, and lets deterministic code turn the answers into moves.

  • typesafe-jev-drone-demo - Simulation: Three.js drone simulator with a Python backend where Jev drives the navigation decisions.

  • typesafe-playground - Interactive playground: small Jev experiments that put the decision on screen, from routing a support message to steering a car in a 3D world.

  • PlayJev - Gaming: open 0.8B vision-language model that reads one 448 px game frame, returns a probability over the moves the game lists in a single forward pass with no generated text, and hands its low-confidence steps to a search program, across ten browser games.

  • jev-plays-pokemon-red - Gaming: Pokemon Red on PyBoy where deterministic code owns the route and arithmetic, Jev picks only at branches, and every battle turn's faint prediction is scored by Brier against RAM state.

  • jevpilot - High-Frequency / Games: A browser driving simulator where Jev chooses among locally generated paths and speeds.

  • jevk5 - High-Frequency / Games: An open model that answers the same typed questions Jev answers — yes/no, choice, score — in one forward pass with zero generated tokens (~13 ms on an H100), plus a head-to-head harness that puts Jev and the open model on the identical board and question so the two can be compared directly. It serves TypeSafe's `/v1/systemone` wire format, so a Jev client can point at it unchanged.

  • laya-vs-jev - High-Frequency / Games: Laya vs Jev: local MLX and hosted AI decisions playing T-Rex side by side, with live metrics and replay recording

  • jev-libero - High-Frequency / Games: Two LIBERO tasks, one control engine. Each demo loads its own JSON task definition. Videos follow simulation time, with decision and physics-preview waiting omitted.

  • RoboJEV - High-Frequency / Games: RoboJEV is a small, inspectable robotics laboratory. JEV receives **structured simulator state, not images**, selects an immediate intent, then selects X/Y/Z directions and a gripper command. A Cartesian controller executes the action using real MuJoCo contacts. Each task has independent physical success checks; model answers cannot declare success.

  • OneVOneJev - High-Frequency / Games: A browser-based 1v1 shooter where Jev reads structured match state and chooses movement, aim, and firing.

  • laya-vs-jev-arena - High-Frequency / Games: Laya (open source, local) vs TypeSafe Jev (API): two AI models race in Snake and fight in a Mortal-Kombat-style arena. Every move is a real model decision.

  • jev-doom-agent - High-Frequency / Games: A browser Doom experiment comparing Jev-controlled players from the same initial state.

  • typesafe-snake - High-Frequency / Games: Snake autoplayer driven by TypeSafe Jev, executing one discrete System One decision per tick with code-enforced legal moves.

  • live-jev - High-Frequency / Games: A browser-based top-down driving simulator using Jev for lane and speed choices, with an optional chat-model comparison.

  • jev-reflex-autonomy-lab - High-Frequency / Games: Multi-drone autonomy lab demonstrating TypeSafe Jev reflex decisions with optional System 2 strategy guidance.

  • jev-askable-arm - High-Frequency / Games: Uses Jev to chain predefined skills for English-language goals in a ManiSkill robot-arm simulation.

  • jevscape - High-Frequency / Games: A RuneBench extension using Jev and a bounded rs-sdk action catalog for RuneScape tasks.

  • jevtown - High-Frequency / Games: A check costs from half a cent (a text that dies in the first wave) to ten cents (one that reaches all 10,000), and takes from 3 seconds to a minute. The interface comes in Ukrainian and English, and so do the personas: a text is read by the crowd that speaks its language, 10,000 Ukrainians or 10,000 English speakers, so there is nothing to choose.

  • heist-one - High-Frequency / Games: Observable browser stealth game where Jev makes typed guard judgments while deterministic code owns the physics world.

  • jev_deep_rl - High-Frequency / Games: This project evaluates a fixed model. It records rewards and decisions without training or updating model weights. A seeded random policy provides a local baseline.

  • JevBird - High-Frequency / Games: A Python Flappy Bird game where code simulates candidate routes and Jev picks one.

  • doom-jev - High-Frequency / Games: A ViZDoom Agent that uses Jev to choose movement, targets and firing from structured game state.

  • jev-little-airways - High-Frequency / Games: An island-airport simulator using Jev for routes, yielding, emergency broadcasts and landing order.

  • soupbase - High-Frequency / Games: Soupbase is a bilingual Chinese-English Turtle Soup game where Jev judges player questions and reconstructions, and the app checks structured Choice results and confidence to decide clearance.

  • jev_vampire_survivors - High-Frequency / Games: TypeSafe's Jev model plays Vampire Survivors on Steam: BepInEx plugin + Python brain + live decision dashboard. Native Linux only.

  • jev-robotics-demo - High-Frequency / Games: A MuJoCo arm demo where local code proposes candidate moves and Jev chooses the target, grasp or release, and completion.

  • jev-arena-nanojev - High-Frequency / Games: Jev Arena is a fully local grid tactical game arena where NanoJev, rule agents and search algorithms make per-step move, attack, shoot, heal, dash and environment-interaction decisions across multiple levels with a Chinese Pygame interface.

  • jev-flappy-bird - High-Frequency / Games: A live demo of TypeSafe's Jev model playing Flappy Bird, one flap-or-wait decision at a time.

  • jev-gamepilot - High-Frequency / Games: **Jev-GamePilot** is a universal autonomous AI gaming agent powered by **Laya (local sub-30ms System One inference)** and **TypeSafe's Jev System One** (`Choice`, `Score`, `Noul`). It captures real-time gameplay at 60+ FPS, fuses instant local reflexes with high-level strategic reasoning, and executes physical hardware inputs across Windows PC games and connected Android phones.

  • jev-gpt - High-Frequency / Games: Cascaded Choice questions that make Jev pick the next word from a word tree instead of generating text.

  • jev_fsd - High-Frequency / Games: **An AI model drives a car through a real city, and you can watch every decision it makes.**

  • jev-market-reflex - High-Frequency / Games: Fast typed AI decisions on live crypto markets using TypeSafe AI Jev.

  • jev-play-ping-pong - High-Frequency / Games: Uses Jev to choose serve direction, return angle and pace in a browser table-tennis game.

  • jev-rl - High-Frequency / Games: JEV Reinforcement Learning: four classic games trained with JEV-powered rewards, reproducible experiments and checkpoint replays.

  • jevarena - High-Frequency / Games: Two Jev Agents play Snake in side-by-side browser panes with visible per-step choices.

  • snake-jev - High-Frequency / Games: Real-time Snake game driven by parallel Jev assessments, deciding optimal turns in a single API call per tick.

  • tsai-civ2 - High-Frequency / Games: An experimental harness where TypeSafe Jev plays classic Civilization II in a browser, computing live action probability distributions.

  • jev-broadcast-lab - High-Frequency / Games: A Jev experiment workbench centered on chess, with additional classification and matching exercises.

  • jev-chess - High-Frequency / Games: Chess moves, evaluations, persona opponents, and game classification using TypeSafe AI System One models. Resolves natural language move intents into legal moves, evaluates positional sharpness and king risk in parallel, and powers historical persona opponents (Tal, Capablanca, Petrosian).

  • jev-flappy-bird - High-Frequency / Games: Jev learns to play flappy-bird game with physics based context and without it

  • jevTrader - High-Frequency / Games: A High Frecuncy Trader made in Rust using Jev as a decision maker.

  • mk-jev-fly-brain - High-Frequency / Games: Compares a fly-connectome spiking simulation, Jev and rule policies in the mk.js fighting game.

  • can-jev-bayes - High-Frequency / Games: How well can Jev make sequential decisions under uncertainty, and how can Bayesian methods help it learn and act more effectively?

  • jev-claim-vs-measured - High-Frequency / Games: A post with ~400k views says TypeSafe's **Jev** is the fastest AI model ever built for trading, makes calibrated buy/sell decisions in under 100 ms, and shows how to build an HFT system on it. The article behind it contains no backtest, no P&L and no hit rate. So I ran the tests: on the raw tape at one decision per second, and at 15–60 minute horizons. It cost **$0.97** of API credit. Everything needed to check me is in this repository.

  • jev-clash-royale-test - High-Frequency / Games: A Clash Royale-style sandbox whose Jev bot decides play-or-hold, card, lane, and depth in one System One call.

  • jev-experiments - High-Frequency / Games: Uses Jev to play Chrome Dino and a local shooter arena while Python executes structured decisions.

  • jev-factorio-agent - High-Frequency / Games: Jev picks what, code owns how - a System One Factorio agent driven by TypeSafe's Jev on FLE

  • jev-practice-speed - High-Frequency / Games: A WebGL demo where you play the card game Speed against a CPU whose brain is TypeSafe AI's Jev. The whole point of the app is to measure and show Jev's decision speed and decision accuracy in real time.

  • jev-synthetic-survey - High-Frequency / Games: New to synthetic survey respondents? [Start here](#new-to-this-start-here). For the raw runs, the scored reports and the code, see [where to go](#where-to-go).

  • jev-table-tennis - High-Frequency / Games: Table tennis vs. TypeSafe's Jev (System One). Every paddle move on the right is a live model decision — no local prediction, just a lookup table and a servo.

  • typesafe-jev-decision-studio - High-Frequency / Games: Fast, calibrated System One decision platform powered by TypeSafe Jev via OpenRouter. Sub-second logprob scoring, transfer curves, zero hallucinations.

  • typesafe-jev-traffic-demo - High-Frequency / Games: This is a **simulation**. It is not connected to, and cannot control, any real traffic signal — Hong Kong's Transport Department publishes no write API for that, only a read-only feed of sensor data. Everything downstream of that feed (the phase timing, the amber/all-red clearance, the safety limits) runs entirely in this process's own memory.

  • [jev-torneo-animales](https:

Truncated — view the full README on GitHub.

awesome
awesome-ai
awesome-jev
awesome-list
awesome-lists
awesome-readme
awesome-resources
calibrated-probabilities
jev
jev-ai
jev-alternative
jev-api
jevbench
jev-model
laya
llm
llm-evaluation
llm-inference
llm-tools
vmodal

Significant stargazers

Leo Chiu

65 followers · starred Sep 2026

v-modal/awesome-jev-tools

A curated list of tools built for Jev — TypeSafe AI's System One model for typed decisions.

754

96 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

List of tools around jev in production setup (r/LocalLLM)

Hello, Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. So, we gathered some tools/frameworks around jev: [awesome jev](https://github.com/v-modal/awesome-jev-tools) This list below answers two practical questions quickly: \- Where is Jev…

2

Oct 6, 2026

README

awesome-jev-tools

A curated awesome list of public projects and practices built on Jev, TypeSafe AI's System One model for typed decisions.

This README is the homepage aggregate of the current category files, so the latest accepted entries are visible here without drilling into subpages.

A curated list of public projects and developer patterns built on Jev, TypeSafe AI's System One model for typed decisions.

What is Jev? Jev is not a chat model. It does not write text or hold conversations.

Instead, it takes unstructured state alongside a typed question and returns a typed decision—such as a choice, a score, or a boolean—accompanied by a confidence rating.By eliminating token-by token decoding, Jev acts as a fast, low-latency decision layer directly inside software.

Developers use it to handle classification, infrastructure routing, rubric scoring, verification gates, and autonomous agent guardrails.Goal of this ListMost discussions about Jev are scattered across launch threads, social media, and one-off prototypes.

This repository centralizes those pieces to answer two practical questions for developers:

  • Production Validation: Where is Jev actively making real decisions in live production workflows?
  • Transferable Patterns: Which decision architectures can be cleanly copied and applied across different industries?

Goal of this list

Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. This list answers two practical questions quickly:

  • Where is Jev already making real decisions in production workflows?
  • Which decision patterns transfer across industries?

Inclusion criteria

We do not include:

  • Generic classifiers, routers, or research agents that merely resemble the pattern without using Jev.
  • Pure theory or opinion without a concrete practice.
  • Launch-hype commentary with no working artifact or reproducible result.
  • Long write-ups inside the list itself.
  • Sources that are private, inaccessible, or too vague to classify.

Curation is not endorsement

Inclusion means one thing: the entry satisfies the inclusion rules above. It is not a quality review, a security audit, or a recommendation. We do not verify that a project compiles, that its tests pass, that its published numbers reproduce, or that its license permits your use.

This matters most for projects that arrive in bulk. When one author releases several repositories on the same day, they commonly share a single scaffold — the same AGENTS.md, CLAUDE.md, STATE.md, and CHANGELOG.md — land in one or two commits each, and may ship considerably more prose than code. Such projects can be entirely legitimate; they are simply unproven. Treat them as leads, not as validated tools.

Before adopting an entry, check it yourself:

CheckWhy it matters
Does the code actually call the Jev API?An entry can read well on a README alone. Look for a real request carrying typed questions, and a parsed answer coming back.
Is there a runnable check?A test, an example with expected output, or a public demo. No check means no evidence that it works.
Do the numbers have a source?Any accuracy, latency, cost, or volume figure should be traceable to the linked page. We strip claims we cannot verify, but the project page itself may still carry them.
How much of the repository is code?Some projects are mostly prompt documents. That can be legitimate — just know which one you are getting.
Is there a license?A few entries have none, which limits reuse and redistribution.

Found something wrong? Open an issue or a pull request — removal is as valid a contribution as addition. Rules for AI-assisted work, project depth, and submission rate live in CONTRIBUTING.md.

Current coverage

Each entry lives in exactly one category. When a project could fit multiple categories, we choose the one closest to its direct application domain. You can browse the list by category below.

Open categories still being seeded

  • Scientific Pipelines — 0 entries

Full list

Classification & Routing

Source file: categories/classification-routing.md

  • JEV Book Tags - Library cataloguing: a calibre plugin asks Jev Noul questions about book genres and subjects, applies configurable per-tag probability thresholds, and preserves existing tags while leaving uncertain results for review.

  • Notra - Marketing analytics: production GEO platform whose NOTRA_JEV_CLASSIFIERS flag routes brand-visibility classifiers off an LLM and onto Jev Boolean decisions at a 0.5 threshold, targeting 300 ms p50.

  • jev-router - Developer tooling: routes Claude Code tasks to the cheapest capable model by asking Jev to choose among candidates.

  • Codex Jev Router - Coding agents: asks Jev Choice and Noul questions to select Codex subagent model and reasoning tiers, with confidence gates and a Sol fallback.

  • jev-router (prismhq) - LLM infrastructure: open-source LiteLLM-based router where a Jev decision picks which model serves each request.

  • pi-jev-router - Coding agents: adds automatic per-request model routing to the Pi coding agent through Jev decisions on Vercel AI Gateway.

  • jcm-router - Coding agents: local proxy that picks the Claude model and reasoning effort per message with a Jev decision while leaving the cached main chat untouched.

  • Switchboard - Developer tooling: assesses a new Claude Code or Codex conversation with Jev, then local confidence policy selects and pins its model and reasoning effort through follow-ups, tool calls, and resume to avoid unnecessary prompt-cache disruption.

  • jev-agent-skill-router - Agent infrastructure: routes agent skill selection through typed, confidence-aware Jev decisions so weak matches are declined instead of guessed.

  • typesafe-jev CV screener - Recruiting: screens a folder of CVs with Jev typed judgments against an editable policy, re-scoring candidates for free when the policy changes.

  • Jev email intent workflow - Back-office automation: async LangGraph workflow gets a typed Jev Choice (invoice or general) and routes each inbound email to the matching handler.

  • unclutter - Browser tooling: WXT extension where Jev decides per page element whether it is clutter, removing it under reusable template rules.

  • typesafe-adblock - Browser tooling: Chrome extension that asks Jev whether each DOM element is an ad, turning ad blocking into a stream of per-element typed questions.

  • DiffJury - Code review: routes each pull request by risk with Jev before a human reviewer is assigned, doubling as a review coach.

  • HA-Jev - Smart home: Home Assistant integration that answers questions about the house as a probability, a choice, or a score.

  • secondlayer - Fault triage: self-hosted Stacks data service whose Slack gate and fault-triage paths both run on Jev decisions.

  • new-api-typesafe-plugin - LLM gateway: adds a native /v1/systemone endpoint to new-api so typed decisions sit behind the same gateway as chat models.

  • duet-agent - Agent harness: keeps a Jev-backed routing table for deciding which model should serve a request.

  • json-render - Generative UI: Vercel Labs' UI framework uses Jev in its compose path to pick which components and actions a rendered interface should contain.

  • omo-jevlike-router - Skill routing: shrinks the skill catalog in a system prompt with one forward pass over a frozen Qwen, routing each request Jev-style.

  • jev-cookbook - Developer education: 15 runnable Node recipes that route support tickets, file documents, categorize bank transactions and label Gmail with Jev Choice and Noul questions, sending low-confidence answers to human review.

  • flue-jev-demo - Agent routing: routes a Flue agent's work with Jev through Cloudflare AI Gateway.

  • sift - Content labelling: Chrome extension that labels every post in an X timeline - substance, humour, chit-chat, promo, junk, or AI-written - with Jev decisions.

  • jev-tree - Classification: A hierarchical selector that asks Jev to traverse branches when a catalog has too many options for one question.

  • litellm - Model Routing: LiteLLM can use Jev to classify requests for its complexity-based model router.

  • oh-my-pi - Model Routing: Oh My Pi includes an optional TypeSafe judgment provider for bounded decisions in coding-agent workflows.

  • jev-model-router - Model Routing: A community Claude Code Templates mod that uses Jev to suggest subagent models and reasoning levels.

  • openchamber - Model Routing: OpenChamber’s optional automatic model router uses Jev to classify a message before selecting a configured model and reasoning level.

  • firstmate - Model Routing: Firstmate can optionally use Jev to match task briefs to dispatch rules before local policy chooses an Agent configuration.

  • hermes-jev-skills - Model Routing: Jev-powered model routing, memory, compaction, skill selection, computer and browser use for Hermes agents (also Claude Code and Codex)

  • vexjoy-agent - Model Routing: An optional Jev routing path that matches VexJoy requests to specialist Agents, skills and workflows.

  • WrongStack - Model Routing: An optional Jev dispatch classifier for choosing among WrongStack specialist Agents.

  • jev-codex-router - Model Routing: Uses Jev to classify each Codex turn, then applies local rules to choose the model, reasoning effort, and speed mode.

  • JevRouter - Model Routing: Routes among models, subagents, skills, MCP tools and CLIs using a shared candidate set.

  • grok-bot-jev - Model Routing: Connect TypeSafe Jev to Grok Bot as a cheap decision layer - usage gates, skill template, examples

  • loki - Model Routing: Loki optionally adds Jev typed-judgment tools and routes a new session to a model within the selected gateway.

  • JevLoop - Model Routing: The agent loop where decisions don't cost a large language model call. Zero deps, runs offline, no API key needed.

  • muse-jev-playbook - Model Routing: Jev decision layer for Muse: a fast, cheap TypeSafe AI gate before expensive agent work — confidence policy, recipes, reference router, honest measurement.

  • sabi - Model Routing: Adaptive inference scheduling for AI agents — per-round model, effort and provider routing for coding harnesses: a Command Code mod or a local OpenAI-compatible proxy.

  • typesafe-skill-router - Model Routing: An opt-in Hermes Agent plugin that asks Jev to suggest one relevant skill before a model call.

  • dejevu - Model Routing: Jev? Déjà vu. Browser agents that run on instinct, no Jev needed. One look at the page, one call to any open model, one action. Faster than the Jev demo on Google Flights.

  • jev-router - Model Routing: Cost-aware LLM router that picks the cheapest model capable of handling a query, using TypeSafe's Jev for fast classification instead of an LLM call.

  • laya-jev-lab - Model Routing: Independent measurements of typed-decision models: Jev (TypeSafe API) vs Laya (open weights), and a local-first cascade that matches Jev's accuracy at 1.8x the speed

  • Jev-Auto-Router - Model Routing: Jev Auto Router (Jev Router): experimental per-call GPT model routing for Codex via TypeSafe Jev and a local Responses proxy, with independent task verification.

  • tool-prune - Model Routing: Calibrated tool selection and schema pruning for AI agents. Dual-engine: zero-dependency offline TurboQuant or TypeSafe System One (Jev). Prunes candidate MCP tools and schemas down to the relevant set before calling LLMs to eliminate hallucinations and save tokens.

  • jev-for-all - Model Routing: Jev for every agentic development workflow — the System One decision model wired into whatever harness an agent codes in: OpenCode today, Claude Code and Hermes adapters next.

  • Jev-Model-Router-Claude-Code - Model Routing: Begleitmaterial zum Video „Jev + Claude Code: 3 Use Cases".

  • jev-opus - Model Routing: Claude Opus 5.5 with the effort level re-decided every step by the TypeSafe Jev reflex — without breaking the prompt cache. CLI + Claude Code plugin.

  • jev-pilot - Model Routing: Let Jev steer Claude Code: the right reasoning effort, subagent model and skill for every prompt. A Claude Code plugin powered by TypeSafe's Jev (OpenRouter / TypeSafe).

  • jev-claw - Model Routing: Typed model routing for OpenClaw agents, powered by TypeSafe Jev

  • jev-model-router - Model Routing: Model router for Claude Code using Jev

  • jev-router - Model Routing: Pass-through model router for Claude Code and Codex CLI that picks a model tier per human turn with Jev, TypeSafe AI's decision model

  • jev-smart-router - Model Routing: JEV Smart Router — a Databricks App that uses TypeSafe JEV to pick which model answers each message, then runs inference on the chosen Databricks Foundation Model API endpoint.

  • opencode-jev-orchestrator - Model Routing: An OpenCode orchestrator that keeps a cheap sticky parent model and, when Jev flags a hard turn, escalates through a child subagent.

  • tiershift - Model Routing: Policy-driven model routing framework routing every LLM call to the cheapest capable tier in ~180 ms via TypeSafe Jev.

  • todo-jev - Model Routing: A task-routing experiment combining skill conditions and environment checks to suggest rules, skills or a large model.

  • chat2jev - Model Routing: Convert OpenAI-compatible Chat Completions requests into TypeSafe System One (Jev) **State / Questions**, compare generated text with structured judgments, and publish reusable question sets as proxy routes.

  • Janus - Model Routing: Framework for measuring when to employ Jev versus generative LLMs on proprietary datasets, routing queries based on measured benchmarks.

  • jev-codex-model-and-effort-router - Model Routing: Copy and paste this into your coding agent:

  • jev-codex-pilot - Model Routing: A Codex overlay incorporating JEV to make the best decisions regarding model selection and depth of reasoning. All while automating the process via an automated Kanban system.

  • jev-route - Model Routing: **Run it. Log it. Distill it. Own it.**

  • jev-routing-experiment - Model Routing: Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena

  • codex-jev-native-router - Model Routing: Experimental native Codex Desktop and CLI model routing with Jev and a configurable allowlist

  • hermes-jev - Model Routing: Jev decision sidekick for Hermes Agent — TypeSafe and Cloudflare, explicit tools and official skill

  • jev-agent-hooks - Model Routing: TypeSafe Jev hooks for Claude Code, Codex and pi: per-turn skill suggestion and subagent model routing

  • jev-claude-router - Model Routing: Model router for Claude Code using Jev

  • jev-model-router - Model Routing: Cost-optimized OpenRouter model router using TypeSafe's Jev, with a live full-catalog scorer instead of a hardcoded model list

  • jev-router-playground - Model Routing: A model-routing playground where Jev picks a candidate and users compare the resulting answers.

  • jevbus - Model Routing: A streaming event bus whose routing, subscription and consumption are decided by a probabilistic judge. The reference judge is TypeSafe AI's Jev (System One) model: send it a payload and a set of typed questions, get back calibrated probabilities instead of prose.

  • openclaw-jev-plugin - Model Routing: A silenced message never reaches the language model, so it costs one Jev call and no model tokens. Direct messages always get an answer unless you choose to gate them too.

  • stuntdouble - Model Routing: Drop-in /v1/systemone proxy that shadows Jev with local decision models (Kev, Laya) and reports whether you can swap

  • hermes-jev-router - Model Routing: Hermes Agent plugin: TypeSafe Jev model routing + trim-then-compress

  • jev-lab - Model Routing: Open lab: Jev (TypeSafe System One) routing in front of Claude Code - measured bugs, patch, and a hard fallback with alerts

  • jev-research - Model Routing: A Jev and Herdr integration guide with a prototype for routing tasks to different Agents.

  • JudgeJev - Model Routing: The public site replays **17 real, recorded Jev evaluations** of fictional support cases. It starts with a false shipping guarantee that scored **15.1%**, alongside a correct control that scored **95.5%**. You can inspect raw requests and responses, adjust decision rules, and see why a five-pair candidate fails a release gate.

  • omp-plugin-jev-router - Model Routing: Route Oh My Pi prompts between simple and advanced models with TypeSafe AI's Jev classifier.

  • jev-cc-codex-router - Model Routing: Per-turn model routing proxy for Codex: asks Jev which tier each task needs, rewrites the model, retries flaky upstream errors.

  • jev-decision-gateway - Model Routing: A gateway that asks Jev continue / tool / verify questions and invokes a generative LLM only when policy says generation is needed.

  • jev-demo - Model Routing: A customer-service routing demo that batches Jev questions before following the resulting route.

  • jev-gateway - Model Routing: Session-aware OpenAI-compatible model-routing gateway powered by JEV

  • Diffusion Jev - Visual classification: independent Jev-style DiffusionGemma/SGLang server that selects doodle and flower labels from image pixels with typed Choice questions and displays candidate scores in a drawing playground, with public evaluation artifacts and uncalibrated probabilities.

  • jev-logtriage - On-call operations: batches collapsed Loki logs into one Jev call of Noul, Score, and Choice questions, then maps answers in code to suppress, watch, review, notify, or page, with low confidence going to review and nothing executed.

  • DocJev - Document pipelines: LlamaIndex's open-source library that classifies a document against natural-language category rules or finds the boundaries between sub-documents, with swappable OCR backends (liteparse or LlamaParse) and a benchmark harness whose 40-document pilot classified 40/40 originals correctly at about 182 ms Jev decision p50.

  • jev-fit - Developer tooling: hosted fit checker that sends a pasted software idea and a fixed typed rubric to Jev in one call, where a Choice picks plain code, Jev or a reasoning LLM behind a Noul gate for non-tasks, code vetoes Jev when the idea needs images, and low confidence returns "not sure"; closed source, free page and API.

  • Jev-Mail - Email productivity: runs a 24/7 Gmail classifier on user-owned Google Apps Script where Jev scores urgency, importance, and category, routing uncertain or suspicious mail to Review without a local daemon.

  • AI-decision-maker - Data cleaning: asks Jev Choice questions to classify CSV columns into a 13-code type vocabulary and each dataset into one of six scenes, then executes every write locally; measured Jev at 6.6–12.7× an LLM's token cost on this task because the output is already one character while per-question criteria repeat.

  • hearth-jev-rental-search - Housing search: autonomous multi-source rental search where Jev decides which listings match the criteria.

  • pi-jev-skill-picker agent: Pi - Coding agents: ranks the Pi agent's installed skills against the current task with Jev before any of them run.

  • Feed Lens type: extension - Social media: uses Jev Noul judgments against per-platform, user-defined topic and expression labels to annotate Weibo, Threads and X posts directly in a Chrome extension.

  • jev-table-import-mapper - Data import: maps an uploaded CSV's columns onto a destination table with a strict deterministic name-equality pass, then one Jev Noul per remaining (source, destination) pair plus a guard Noul per incoming column, mapping 10 of 10 columns of a 23-column export at 253 questions in one call, 915 ms, $0.0012, unmapped fields left visible above a 0.75 threshold rather than guessed.

  • Jevidence - Developer education: Python sandbox asks Jev Choice and Noul questions about issue category and reproduction steps in opt-in live mode, then applies confidence and reproduction gates to propose a queue or review fallback without assigning the issue, with synthetic offline fixtures and policy tests.

  • langchain-skill-router - Agent infrastructure: per-turn skill routing for LangChain deepagents, where Jev ranks the SKILL.md catalog against the request and the recent conversation and verifies the top candidates, so only the picked skill's instructions reach the prompt; the judge is a protocol that a self-hosted model or static rules can implement instead.

  • jev-rental - Consumer rental: sorts every claim in a rental listing into verify-on-site / demand-evidence / high-risk-pitch buckets to build a pre-viewing checklist with code-templated questions; 50-sample calibration reports 0.910 gated accuracy and 0/10 injection flips.

  • jev-resume-disqualifier - Recruiting: knocks a resume out of a pipeline in under 25 ms by asking Jev the disqualifying question first, so only survivors reach a full evaluation.

  • Jev-IOT - Smart Utilities & Telecommunications: Ultra-low-cost, non-autoregressive AI telemetry classifier enabling sub-150ms anomaly triage and autonomic remediation across 10M+ smart meters for under $35/month.

  • AgentScope - Multi-agent platforms: multi-agent platform by Alibaba implementing native TypeSafe Jev classification models for binary, choice, and score routing across agent pipelines.

  • inbox-zero - Email productivity: open-source AI email assistant that uses TypeSafe Jev System One decision models to classify incoming email intent and triage action items.

  • SiYuan - Knowledge management: privacy-first personal knowledge management system featuring native Jev decision model integration for high-speed document classification, flashcard intent categorization, and automated tag routing.

  • Paca type: self-hosted - Project management: self-hosted open-source Jira alternative that auto-assigns tasks with a Jev Choice over member descriptions, fills blank task fields with Choice and Score questions, and routes automation workflows on a Choice/Score/Noul condition node, applying answers only at 0.6 confidence or above and otherwise leaving the task unassigned or taking the Else branch.

  • Qualm - Digital wellbeing: macOS menu bar app that reads the screen as text through the Accessibility API and asks Jev (or Kev, its local open-source counterpart) one Choice per user rule plus a Noul on whether the page is a payment, login or banking screen, stepping in with a pop-up only when a rule's probability clears its threshold and never on sensitive pages; on 119 trial pages with Kev, the short-video, feed, livestream and video rules had precision 1.00.

  • Auto-optimizing Jev: half the errors, 1/7 the cost - Text classification: asks Jev a Choice over the readings of a Chinese polyphonic character while the model stays fixed and only the harness around it is optimised, ending at half the errors for a seventh of the cost.

  • spending-effort-with-jev agent: Claude Code type: plugin - Coding agents: Claude Code plugin whose UserPromptSubmit hook asks Jev a Choice over /effort levels (low / medium / high / max / unclear) plus a Noul on whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set.

  • tab-jev type: library - Tabular prediction: asks Jev a Noul on the target plus Score rubrics about each row's text, turns every option's probability into a column next to the row's numeric fields, and lets a tabular foundation model such as TabPFN learn from the labeled rows in context, reaching 0.745 AUC at 256 labels on Kickstarter funding against 0.682 for Jev alone with calibration.

  • tinystruct-typesafe-sdk - SDK: TypeSafe Jev integration library for building type-safe classification and routing decisions with structured outputs.

  • sortwell agent: Claude Code type: plugin - Personal inbox: MCP server and Claude Code plugin that files each captured note, link or meeting line with one Jev request of Choice questions for kind, project and next action plus a Noul for duplicates, routing to a project only at 0.45 or above and marking a duplicate only at 0.70 with a specific matching item, while the text itself is stored verbatim in append-only local files.

  • IntentSQL - Natural-language SQL: turns a question about a SQLite database into a sequence of small Jev decisions instead of one generated query, released as an experiment alongside its decision lab.

  • TypeSafe Conversation - Home automation: a Home Assistant voice agent built on Jev.

  • Jevvie type: library - Web companions: a page offers its actions as WebMCP tools and one Jev Choice picks the action a visitor's request means, with a Choice per argument asked alongside, asking back when the top two options are close and gating unprompted tips with a Noul (source).

  • Gut Check - Smart home: Home Assistant integration whose eight install checks ask Jev a Score on each pending update's release notes and a Choice per item elsewhere, such as whether an unavailable entity is expected, worth fixing or safe to remove; answers below 0.5 confidence change nothing, and the rest that need action become Repairs cards the user must confirm.

Adaptive & Realtime UI

  • DWIM - Desktop productivity: a macOS command palette that reads the frontmost app's menu tree through the accessibility API, asks Jev one Noul per menu item against the user's plain-language request, and presses the top match when it clears a probability threshold, falling back to a ranked list otherwise and never auto-running destructive items.

  • shapeshift - Input: one text box that morphs into the right UI as you type, asking Jev which control the sentence calls for, and running offline.

  • Jevcast - Desktop productivity: native macOS launcher and window manager that uses Jev to match natural-language window and action commands to known application workflows with local response caching.

Verification & Guardrails

Source file: categories/verification-guardrails.md

  • GeekLink Jev Subtitle Translator - Media localization: reviews source–translation SRT pairs with one Jev Noul decision per cue and flags suspected omissions or meaning changes for human review.

  • is-malicious - Software supply-chain security: asks Jev Noul checks about source and build files, escalates suspicious chunks for a second pass, and returns implicated files and lines before execution.

  • jev-review - Software engineering: staged code-review workflow and local dashboard where Jev gates each review stage before a change advances.

  • pi-jev - Agent safety: adds a measured tool-call gate to the Pi coding agent so risky calls are checked by Jev before execution.

  • OpenWork - Engineering workflow: wires Jev into its eval testkit as a verification judge so agent-produced work is gated by typed verdicts rather than a text model.

  • jev-guard - Agent security: prompt-injection and dangerous-action guard for Claude Code, Codex, Pi, and ACP agents, with Jev deciding what to block.

  • Foreman - Software factory: sits above Codex workers and has Jev independently judge whether an implementation is complete, its tests sufficient, or a human is needed.

  • stanley-code - Coding agents: bounded Jev workflows that keep agent judgments typed instead of free-form.

  • opencompany - Agent workspace: runs its approval review through Jev so workspace actions are gated by a typed decision.

  • jev-git - Developer tooling: sub-second Git pre-commit & pre-push reflex gate that screens staged diffs for secrets and destructive commands using Jev.

  • pi-heed - Runtime constraints: checks every side-effecting tool call from the Pi agent against what the user actually asked for.

  • Hunch - Code review: plain-English rules that Jev checks code against, locally or on every pull request, with Jev picking one label per finding.

  • Abide - Agent supervision: reads every edit a coding agent makes and has Jev flag rule violations, with the project reporting that an independent reviewer confirmed 10 of the 39 flagged edits and 11 of the 15 flagged turns.

  • fx - Coding agent: ships a typesafe_permission_reviewer builtin so the agent's permission decisions run through Jev rather than an LLM call.

  • Sniff Test - Writing: prose linter that asks Jev ten Boolean questions per paragraph (stacked hedges, restating closers, not-X-but-Y turns, naked cost figures) at a 0.7 threshold; CLI, pre-commit hook, GitHub Action and Claude Code skill; measured 182 ms median and 1 of 54 clean paragraphs flagged against 37 for Haiku 4.5.

  • jev-pref - Code review: turns the preferences in a project's AGENTS.md into jev-pref.json rules that Jev checks against each diff hunk, staged file set, or pull request, returning fix_now or advisory findings to the coding agent and a nonzero exit code on blocking ones.

  • jev-axi - Agent safety: PreToolUse gate for Claude Code and Codex that has Jev score each shell command for destructiveness, exfiltration, remote code execution, and security weakening, deciding routine commands locally so nothing is sent for them, and scoring 44/44 on the 44 labeled tool calls in its repository.

  • pi-verdict - Agent safety: Pi permission gate where Jev answers one Choice (allow/ask/deny) per gray-zone tool call — deterministic rules settle clear cases first, deny blocks, ask escalates to a human confirm, and errors or timeouts deny; Jev is an optional backend, OpenRouter-only and experimental.

  • jev-commit - Developer tooling: pre-commit hook where one Jev call judges whether the commit message matches the staged diff, flags debug leftovers and unmentioned work, and blocks only on a detected credential.

  • Blink - Code review: CLI that coding agents run after every change, with Jev checking the diff near-instantly in place of an LLM reviewer.

  • hermes-jev-approvals - Agent approvals: proof of concept that puts Jev in front of Hermes Agent's command approvals, reporting 8.7x faster decisions and 4.4x fewer prompts to the user.

  • jev-engineering - Agent safety: gates coding-agent tool calls with deterministic rules first and one typed Jev call second, then publishes a rerunnable 300-call injection test showing what the gate catches and what walks past it.

  • Cribrix - Retrieval / RAG: filters retrieved chunks with a Jev Score plus Noul checks for answer evidence and prompt injection, then withholds any draft whose claims fail a batched per-claim Noul or cite numbers absent from the sources; on its replayed 62-question golden set it answered 0 of 22 unanswerable questions, against 4 of 22 for naive top-5 RAG.

  • agentgateway - Security & Guardrails: A Jev guardrail example in Agentgateway using a webhook to inspect model requests and responses.

  • Agent - Security & Guardrails: An optional Jev command-risk advisor inside a native macOS Agent, with a TypeSafeKit client.

  • interlinked-cli - Security & Guardrails: Interlinked adds optional Jev judgments and evidence checks to local coding-agent checks.

  • pi-warden - Security & Guardrails: Adds checks for project rules, out-of-scope actions, repeated failures, and completion claims to Pi Agents.

  • building-with-typesafe-jev - Security & Guardrails: Unofficial skill that teaches coding agents to build with TypeSafe AI's Jev: typed decisions, calibrated confidence, and prior art from 150+ community projects.

  • jevals - Security & Guardrails: Agent evals and guardrails as Jev decisions: one request per trace, a fraction of a cent, fast enough for the agent loop. Runs locally with Kev or Laya.

  • captaincore - Security & Guardrails: Jev commands in the WordPress toolkit CaptainCore answer structured questions and prioritize malware scanner findings for review.

  • jev-kit - Security & Guardrails: Everything you need to run TypeSafe's Jev with Claude Code: a tool-call guard, tier guard, file search, browser agent, review, belay, compaction and installers.

  • jev-edge - Security & Guardrails: Typed-judgment admission control at the traffic edge: three-layer prompt-injection and abuse filter for nginx/OpenResty, powered by TypeSafe Jev. Fail-open, cached, hot-reloadable.

  • pi-jev-auto-mode - Security & Guardrails: Adds rule checks to Pi commands and file operations, then uses Jev to assess cases that need further judgment.

  • jevvy - Security & Guardrails: Jev-powered plugins for coding agents

  • jev-safety-gateway - Security & Guardrails: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-security-scan - Security & Guardrails: Reviews Agent Skills and MCP configurations and source code with local static checks and TypeSafe Jev before installation or execution, reporting file and line evidence, risk categories, model probabilities, and coverage gaps.

  • jev-enforce - Security & Guardrails: 📏 Claude Code plugin that makes Claude follow your CLAUDE.md: every reply and edit checked by TypeSafe Jev ✅

  • jev-engineering - Security & Guardrails: Jev Engineering: Typed Decision Systems for Reliable Agent Workflows. Paper, diagrams, and companion examples by Av1dlive.

  • JevPR - Security & Guardrails: PR Risk review, automated by Jev

  • pi-jev-sentinel - Security & Guardrails: Open-source Pi coding-agent extension that uses Jev to check tool calls before they run, scan files for prompt injection, flag risky replies, and keep secrets out of what it sends.

  • dsh-jev - Security & Guardrails: Jev (System One decision model) plugin suite for DeepSeek Harness (dsh)

  • jev-auto-approve - Security & Guardrails: Jev is a decision model: it answers a typed question with a calibrated probability rather than prose. This action asks it one yes/no question per thing worth being sure about — answered in parallel in a single call — and approves only when every one of them clears your threshold:

  • jev-block-android-ad - Security & Guardrails: An Android notification and SMS filter that applies local OTP rules before asking Jev whether a message is advertising noise.

  • jev-guard - Security & Guardrails: Probability-scored guardrails for Claude Code: deny rule-breaking edits and unasked-for deploys, route your docs into each prompt, and check the final answer against the turn's own evidence.

  • jev-phishing-bench - Security & Guardrails: The signal result above was challenged on three points: no non-AI baseline, selection and evaluation on the same emails, and no equivalent decomposition for the LLM. Three controls were added (`bench/heuristics.py`, `bench/protocol.py`, `run_llm_signals.py`); nothing above was changed. Full tables in `results/report.md`, chart in `results/controls.png`.

  • jev-tool-permissions - Security & Guardrails: Adds tool-call approval and tool-list pruning to the Vercel AI SDK.

  • pi-jev-guard - Security & Guardrails: Check Pi code edits against repository Markdown rules with TypeSafe Jev

  • agi-jev-containment - Security & Guardrails: **Open-source AI agent monitoring, malicious-agent detection, and escalate-only containment** for sandboxed LLM agents. Local HackSpain 2026 stack (AngryRobot dashboard): FastAPI, React/Vite, Neo4j. Classifies a *chain of actions*, not a single tool call. A model never pulls the plug.

  • jev-baselines-eval - Security & Guardrails: Every latency number here is **wall-clock duration of one API call** measured in the client: a timestamp before the request, another after the full response body is read ([`code/run_b1.py:32`](code/run_b1.py), [`code/common.py:50`](code/common.py)). Non-streaming on both sides, so these are completion times, not time-to-first-token.

  • jev-model-tokengate - Security & Guardrails: An OpenAI-compatible proxy that sits between your LLM and your users. It evaluates each sliding window of tokens **while the response is still streaming** and cuts the stream **before** a violating token can reach the screen.

  • open-jev-approvals - Security & Guardrails: Binary approval gate for Codex and Claude Code — every intercepted tool call is reviewed by TypeSafe JEV and composed through a versioned local policy, with scoped authorization.

  • reflex - Security & Guardrails: A coding agent and personal assistant built on the Pi coding agent. Jev checks every tool call, turn and voice transcript, and code decides what happens next: allow, ask or block an action, which model tier to use, and whether a "done" was actually verified.

  • ego-jev-ultrafast - Security & Guardrails: Jev drives your Ego Lite browser: one typed-choice request per step. Single-file, zero-dependency port of browser-use/jev-ultrafast with multi-model benchmarks and extra guardrails. Unofficial.

  • jev-decisions - Security & Guardrails: Jev Decisions Plugin for Hermes (and other AI Agents): tool risk reviews, human approval recommendations, evidence checks, and a local decision journal.

  • jev-guard - Security & Guardrails: High-speed, cross-agent safety gate plugin for **Claude Code**, **Codex CLI**, and **Antigravity**.

  • jev-secret-detection - Security & Guardrails: Benchmark and tool evaluating how well TypeSafe Jev identifies real secret credentials in file snippets.

  • jevshield - Security & Guardrails: Sub-100ms security gate for AI agent tool calls, powered by TypeSafe's Jev (System-1) decision model. Single-request Choice/Noul/Score evaluation, dual-factor blocking matrix, calibrated-confidence routing, fail-closed parsing, zero-config local fallback. LangChain-ready.

  • oc-plugins - Security & Guardrails: The oc-auto-perms plugin in an OpenCode plugin collection uses Jev to check tool intent against natural-language rules.

  • actiongate-jev - Security & Guardrails: Open-source Jev tool-calling authorization gateway for AI agents: deterministic policy, exact-action single-use permits, MCP and HTTP enforcement.

  • antivirus - Security & Guardrails: A file scanner that sends extracted features to Jev for a verdict, a 0–4 severity score, and Noul indicators, then applies local quarantine or review rules.

  • claude-jev-plugin - Security & Guardrails: TypeSafe Jev semantic guardrails for Claude Code

  • dsh-jev-verify - Security & Guardrails: Jev (TypeSafe System One) decision tools + live verification benchmark for DeepSeek Harness: jev_decision (choice/score/noul) and jev_verify, honest by design.

  • grok-jev-guard - Security & Guardrails: `grok-jev-guard` sits immediately before a meaningful Grok Bot tool sequence. It receives a compact description of the pending operation and returns one explicit action:

  • jev-cvss - Security & Guardrails: Scripts that use Jev to select CVSS metrics from vulnerability descriptions, then compute v3.0, v3.1 or v4.0 scores in Python.

  • jev-pii-checker - Security & Guardrails: A CLI that sends text to TypeSafe Jev for PII category Nouls and a sensitivity Score, then locates spans with regex and segmentation.

  • jev-risk-check-provider - Security & Guardrails: x402check — LIVE payer-intent risk checks for x402 agent commerce: typed decisions (TypeSafe Jev), ES256-signed attestations, mainnet USDC settlement. did:web:x402check.xyz

  • opencode-jev-guard - Security & Guardrails: When FarHand is active, the agent's commands run on a remote host through the `farhand_remote_shell` MCP tool instead of `shell`. OpenCode's permission request for an MCP tool carries no arguments, so the plugin takes the command (and its `cwd`) from the `execute.before` hook, which OpenCode runs first.

  • claude-jev-warden - Security & Guardrails: Real-time quality gate and Art Director Warden for Claude Code powered by TypeSafe Jev 1.13 non-autoregressive decision model

  • guardrail-chatbot-jev - Security & Guardrails: It is a library, not a service. You call it, you get a verdict, and your code decides what to do. It runs in **Python and TypeScript**, both reading the same policy file, so the two sides of your stack cannot drift apart. Neither package has a third-party dependency.

  • jev-chrome-extension - Security & Guardrails: 1. Open any website. 2. Click the Jev icon. The side panel opens on **Drive**. 3. Type a goal, e.g. *Search Wikipedia for "espresso" and open the article*, and press **Run**.

  • jev-preflight - Security & Guardrails: A bounded Jev risk check for Claude Code: eight risk axes, one request, one optional reinspection.

  • jev-reasoning-navigator - Security & Guardrails: En lugar de depender de heurísticas matemáticas frágiles o distancias vectoriales locales de coseno, `JEV-Reasoning-Navigator` utiliza **TypeSafe AI (`typesafe-sdk`)** como motor único y autoritativo de decisión cognitiva:

  • jev-test - Security & Guardrails: Prototype: AI-assisted NZQA marking from rubric criteria alone, using TypeSafe Jev for guardrails, criterion scores and confidence-based triage. NOT ENDORSED BY NZQA - CONCEPT ONLY

  • momus-review - Security & Guardrails: Code review on rust and javascript (+languages soon) applications. Following lenses correctness, security, reliability, compatibility and testGap.

  • pkg-gate - Security & Guardrails: Pre-install security gate for npm lifecycle scripts using TypeSafe System One. Evaluates preinstall, install, and postinstall hooks across intent, threat severity, secret access, and remote execution to intercept supply-chain attacks before execution.

  • traffic-guard - Security & Guardrails: High-throughput traffic and attack defense gate for incoming HTTP traffic with zero required dependencies, wire-order header validation, and TypeSafe System One acceleration for bot mitigation, exploit detection, and risk scoring.

  • laya-browser-guard - Security & Guardrails: A passive, privacy-first Chrome security copilot that combines deterministic browser-visible checks with local Laya and official Jev typed decisions. It evaluates redacted evidence from scripts, resources, forms, headers, and DOM signals, then explains investigation priority without attacking the target.

  • Edward - Agent operations: one batched Jev Choice over the cross-turn coding-agent trajectory decides continue, pause, or escalate, with low-confidence verdicts routed to a human while deterministic code keeps dangerous-command blocking, budget caps, and an Ed25519-signed receipt chain.

  • taste-lint - Writing / UI: CLI that uses Jev probabilities on semantic taste checks to catch AI slop in UI, copy, and agent instructions before ship; measurable rules stay local and active findings can fail a run.

  • jev-harness agent: Multi type: cli - Developer tooling: System 1.5 quality gate and token optimizer for AI coding agents that triages test failures in < 500 µs to resolve missing dependencies without frontier LLMs, aborts circular doom loops, and modulates reasoning effort across Python, TypeScript, and Rust.

  • TryJevAI - Scheduling: public Jev playground uses a typed Choice with an explicit Unresolved option to distinguish a mentioned arrival time from an agreed meeting time, showing the returned probabilities and prompting for missing agreement before treating a time as settled.

  • Agent Chaperone - Agent safety: MCP proxy plus hooks that screen a tool call before it runs and a tool result before the agent reads it, with 45 test files behind it.

  • jev-proof - Creator sponsorship: verifies each sponsored ad segment in video subtitles against acceptance rules with one Jev Noul+Choice call per fact while deterministic code keeps the confidence gate; 90-sample calibration reports 0.922 gated accuracy and 0/15 injection flips.

  • jev-fidelity - Editorial QA: asks Jev per fact unit whether an edit preserved the original (preserved / equivalent / drift / lost) behind a 0.70 confidence gate in code, degrading to human review rather than pass; 55-sample calibration on real Wikipedia revision diffs reports 91/92 gated judgments correct and 0/20 injection flips.

  • approval-judge-bridge type: proxy - Agent safety: OpenAI-compatible /v1/chat/completions proxy that gates an agent's shell commands through a calibrated Jev Choice decision with fail-closed semantics.

  • Dub - Link safety: calls typesafe-ai/jev in malicious-link-check.ts before a short link is created, so the URL is gated by a typed verdict rather than a blocklist.

  • JevGate type: cli - Code review: CI and coding-agent gate that parses code locally and asks Jev Noul, Choice and Score questions about one function, file outline, candidate copy pair or test at a time, turns answers at 0.80 into review or consider findings with file and line, fails the build on review, and keeps undecided files as uncertain instead of clearing them.

  • dsh-jev-interceptor agent: DeepSeek Harness type: plugin - Coding agents: DeepSeek Harness plugin where a Jev Choice risk class plus Noul irreversibility, task-match, and injection checks gate every non-read-only tool call (deny confident high-risk, ask ambiguous, delegate the rest), Noul scope and reversibility questions auto-approve clearly-granted calls behind argument-evidence gating, and a per-message Score re-ranks what a referenced session keeps instead of oldest-first dropping — fail-closed to stock behavior, shadow mode with a /jev-stats command, 64 tests.

  • jevci type: cli - Quality gate: asks four typed questions about each change — three Score lenses and one Noul — and blocks a diff, commit message or doc set that falls below the resulting quality score, from the terminal, a pre-commit hook or a GitHub Action.

  • pi-subagent-jev agent: Pi type: plugin - Agent governance: when the Pi main agent dispatches a subagent, evaluates the task text against configurable rule sets in one typed Jev call (per-rule probability questions with below/above thresholds), blocks the dispatch with per-rule reasons on any hit, and fails open to allow on errors.

  • jev-lint agent: Multi - Software engineering: uses Jev Noul judgments and local thresholds to flag team-rule violations as Claude Code and Codex edit, helping agents fix them before code review with configurable rule packs and repository-specific rules.

  • jev-secret-guard agent: Claude Code - Agent security: Claude Code PreToolUse hook that blocks known key formats locally and sends unknown high-entropy strings to Jev only in masked form for a Noul on whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable; 6 of 6 secrets and 0 of 6 benign strings were blocked in its published calibration.

  • Perch agent: Multi type: cli - Code linting: semantic code linter that asks Jev about each method with its callers and callees in view, a Noul for whether it has a bug, a Choice for which kind and which line, and a Score for severity, plus language-filtered CWE Noul checks and custom rules written as sentences at repository, file or method level, failing CI on any answer over its floor.

  • semcheck type: cli - Code review: Go linter whose rules are plain-English questions such as "does this log call write personal data?", asking Jev one Noul for each piece of code a rule applies to and reporting it above the rule's threshold; its two shipped rules were right on 12 of 12 sampled findings in three open-source projects.

  • Skill Scanner - Agent security: Cisco's scanner hunts prompt injection and exfiltration in agent skills, and ships a System One analyzer as a deliberately advisory tier that cannot emit a finding or change a severity.

  • openclaw-jev-leakguard agent: OpenClaw type: plugin - Agent security: OpenClaw plugin that checks every outgoing agent message against where it is going, running local key-format and term rules and then five Jev Noul questions in one call (credential, where credentials are kept, client name, internal infrastructure, confidential business information) through OpenClaw's decisionModel, hosted Jev or a local Kev, and blocking, asking or holding it back by the channel's public, shared or private tier; with Jev it missed 0 of 56 synthetic leaks, 30 of which no regex or term list could see, with 4 false alarms on 57 ordinary messages at 223 ms p50.

Scoring & Ranking

Source file: categories/scoring-ranking.md

  • Clean Code Judge - Code quality: scores every file of a pull request on 31 boolean Clean Code smells plus function size and nesting, then hands the verdicts to a writing model for the review prose.

  • citation-verifier - Academic publishing: checks whether each cited paper actually supports the sentence citing it, with Claude locating the quote, Jev scoring the support, and a human making the final call.

  • jev-bfs - Search tooling: finds link paths between English Wikipedia articles by having Jev rank each page's outgoing links while Python controls the search.

  • Jev Search - Web search: uses Jev Noul judgments on result titles and snippets to rank Search1API results by relevance, with application code merging duplicate URLs and grouping lower-scoring matches separately.

  • pagegrade - Content quality: grades page sections for clarity, writing, and on-page SEO with Jev and returns per-section scores.

  • jev-scout - Developer tooling: sub-second zero-hallucination open-source repo and crate scout using TypeSafe Jev speculative fan-out scoring.

  • jev-seo - Zero-cost, agent-first SEO & Generative Engine Optimization (GEO) search radar CLI suite and MCP server powered by DuckDuckGo and TypeSafe Jev System One.

  • JevSlop - Writing quality: scores public note.com articles on eight Jev Score axes inside a single systemOne request and turns them into a 0-100 Slop Score in ordinary TypeScript.

  • SemanticSpace - Semantic mapping: places phrases in 2D by asking Jev how strongly each one relates to two chosen axis concepts and using those scores as coordinates.

  • Supercov - Code quality for coding agents: Jev answers twelve Noul properties per source file so the agent knows what to fix first.

  • jev.nvim - Developer tooling: Neovim plugin that splits the buffer into functions with Treesitter, scores each against a plain-language question with Jev, and ranks answers by probability in quickfix.

  • jev-reranker - Retrieval and RAG: uses Jev Noul judgments to assess retrieved documents for relevance and usefulness as answer evidence, then sorts results and optionally filters them using a configurable threshold.

  • jev-skip - Media: browser extension that reads the YouTube caption track and scores each segment's sponsor probability on the seek bar before the intro ends, reporting 77% of SponsorBlock's sponsor seconds caught over 23 videos at $0.0008 a video.

  • Refix - Growth: AI that helps your product grow faster on autopilot by running product experiments, SEO, content, and ads.

  • jevsearch - Site search: shadcn/ui ⌘K search block that streams local keyword hits, then sends the top 20 to Jev in one request (a Noul per candidate, a Choice for the best page, a Noul for whether any page answers) to re-order or drop hits, reporting Hit@1 of 83% versus 41% for keyword search alone on 41 labelled queries over the TypeSafe docs (author's benchmark).

  • JevPDF - Document search: browser PDF viewer that extracts each page's lines with pdf.js, asks Jev one Noul per line ("does this line answer the query?") in batches of up to 16 lines sharing the page text as state, and highlights lines ranked by probability as each page returns; only text reaches Jev, through a key-holding proxy.

  • slop-grader - Content quality: CLI tool that grades text files against custom rulesets for AI slop, grammar, and technical doc quality using Jev scores and line-level flags, then guides an AI agent to auto-fix violations.

  • jevseo - SEO and AI-answer visibility: a deterministic crawler extracts every page of a business site, Jev answers a narrow typed Choice/Score/Noul question set per page, and application code turns those probabilities into ranked findings under three confidence bands with the grey zone routed to a needs-a-human pile rather than acted on; it runs locally on one port with no API key through the keyless Zen tier, publishes no search-volume numbers at all because Jev carries no index or volume data, and its source is UNLICENSED (all rights reserved) - unrelated to the other jev-seo entry above.

  • jev-ai-detector type: extension - Writing analysis: Chrome extension which gives readers an instant, uncertainty-aware signal for how strongly selected webpage text resembles AI-generated writing, using Jev inline in Chrome without interrupting reading.

  • Tweet Radar - Social reading: uses Jev Noul to score already-loaded X posts against a reader's goal and profile, then pairwise Choice judgments to rank eligible matches and surface up to three for review.

  • nlgrep - Developer tooling: uses Jev Noul judgments to find code, docs, logs, and text satisfying natural-language conditions, with a configurable probability threshold and ranked file results linked to source lines.

  • hippo-memory - Agent memory: a biologically-inspired memory store whose optional Jev reranker lifts recall R@1 from 0.41 to 0.62 on a private 300-query developer store.

  • MemSearch Jev reranking - Coding-agent memory: an optional Jev reranker asks Noul questions about retrieved Markdown chunks and sorts them by relevance to the query, with bilingual evaluation results.

  • Oko - Developer tooling: local code search for coding agents that shortlists function-level chunks with ripgrep and BM25, asks Jev a Noul relevance question per chunk across three parallel requests, and returns the accepted ones as excerpts through MCP; the cutoff and excerpt selection live in code.

  • grokbot-jev-jobs - Job search: a daily Vercel cron that scores public job postings against one resume with Jev through the Vercel AI Gateway, so only the plausible matches surface.

  • jeff type: cli - Developer tooling: Go CLI whose rank command asks one Jev Score per item per weighted dimension of a YAML spec in a single request and sums weight times score in code to order the items, with noul, choice and score commands that turn a threshold into exit code 10 for shell scripts and CI.

  • OpenViking - Reranking: Volcengine's agent context database ships a Jev rerank client that scores each candidate document with jev-latest against api.typesafe.ai and treats the returned probability as relevance, because TypeSafe exposes no native rerank endpoint.

  • Jev-Code-Reviewer agent: Multi type: cli - Code review: asks Jev for a priority score per changed unit and returns a priorityGap that a local uncertainty policy turns into the order a human should read the hunks in, while OpenAI explains the ones that surface.

  • WorldMonitor - Geopolitical intelligence: real-time global intelligence dashboard using TypeSafe Jev questions to score news headline severity into 5 threat levels and categorize events across 14 conflict, cyber, and infrastructure domains.

Agent Decisions

Source file: categories/agent-decisions.md

  • jev-social - Social media research: uses a Jev Choice at each step to select a concrete socai CLI operation and observed post or profile target on Instagram, TikTok, or LinkedIn, rejecting malformed or low-confidence decisions before execution.

  • Jev Ultrafast - Browser automation: browser-use's ultrafast agent where Jev decides each next action and element to click, calling a language model only when text must be typed.

  • jev-agent-browser - Browser agents: a parent agent delegates bounded tasks to a Jev loop that selects typed browser actions, validates them through agent-browser, and escalates ambiguity or stuck states back to the parent.

  • pi-typesafe-jev - Coding agents: exposes System One judgments as five Pi tools so a model makes narrow semantic judgments while code and users keep control of thresholds, weights, and actions.

  • jev-judgment - Coding agents: agent skill that sends closed coding-agent judgments to Jev so verdicts stay typed, cheap, and comparable across runs.

  • limpet - Coding agents: Stop hook that keeps an agent from finishing too early by judging plain-language completion rules with Jev.

  • robo-harness - Robotics: SO-101 arm workbench where a Jev decision runner picks bounded joint steps from typed candidate actions under a spend budget.

  • dsh-auto-mode - Coding agents: DeepSeek Harness permission preset whose end-prompt step has Jev answer the open questions an agent leaves in its final message, steering them back only when a choice clears 0.6 confidence and an autonomy-safety Noul clears 0.5, and returning the turn to the human otherwise.

  • augustus - Coding agents: agent skill that maps Choice, Score, and Noul onto classical methods so an agent can place typed judgment in software, with a composition algebra, question-design diagnosis, and a validation gate that requires a falsifying experiment.

  • yoshi - Context management: proxy for Claude Code and Codex where Jev judges which conversation history is still needed before pruning.

  • pi-jev (TheoOliveira) - Coding agents: semantic tool routing and typed System One decisions for the Pi coding agent.

  • pi-quiet-ask - Coding agents: gives the Pi agent a quiet Jev decision layer for judgments it would otherwise hand to a chat model.

  • fastbrowse - Browser agents: Jev picks each action from what is on the page while an LLM reads and plans.

  • super-jev - Decision harness: turns a Jev answer into a bounded action instead of leaving the caller to interpret it.

  • jev-superpowers - Systematic software development framework for AI coding agents upgraded with TypeSafe Jev System One typed decisions, zero-hallucination package vetting, and completion gates.

  • Jev Browser - Browser automation: drives a browser with Jev deciding each step, pitched as fast and very cheap next to LLM-driven browsing.

  • pi-fast-jev-compaction - Context management: Pi extension that keeps conversation text verbatim while pruning stale tool history with Jev, falling back to Pi's own summarization only when pruning cannot free enough room.

  • Atomic - Coding agent runtime: ships a first-class Jev structured-output provider so an agent's decisions come back typed, through the same decision resolver as its other providers.

  • fast-jev-compaction - Context management: Claude Code plugin that replaces the compaction summary with Jev decisions, scoring every tool call and result for whether it is still needed instead of summarizing the session.

  • fast-dev-compaction - Context management: Codex port of the Jev-guided compaction idea, restoring context verbatim around a session compaction rather than summarizing it.

  • public-browser - Browser control: lets Claude Code and Cursor drive a real Chrome profile, with a Jev loop deciding the actions, reporting roughly 30% fewer tokens and 25% lower cost.

  • pi-typesafe-router - Coding agents: routes Pi's work through typed Jev decisions.

  • wakegate - Long-running agents: before a sleeping agent's LLM is resumed on a timer or incoming event, Jev answers a Choice (wake, not yet, unrelated) against the agent's own sleep note, and code skips the wakeup only when wake is below 0.2 while always waking on user messages, bare timers, a skip limit, errors, and timeouts; one run passed 21 of 21 hand-written scenarios, which the README calls a smoke test rather than a benchmark.

  • BrowserClaw - Browser automation: Zero-lock, session-preserving Chrome MCP server that couples a local Jev System One semantic micro-loop (chrome_act_toward_goal) with an 85%+ pruned DOM tree (Shadow DOM & iframe pierced), dispatching native CDP events (isTrusted: true) on active logged-in sessions without focus theft.

  • jev-canvas - Multimodal UI: draw on a tldraw canvas by voice while pointing a webcam-tracked finger; on every partial transcript Jev answers eight typed questions (is it a command, is the sentence complete, action, shape, colour, target, place, size) and plain code gates them with thresholds, in English and Ukrainian, 300–550 ms per decision.

  • jev-belay - Coding agents: Claude Code Stop hook that reads the transcript for evidence and spends one four-question Jev call only when files changed with no passing check since, failing open on any error.

  • Jev for Chrome - Browser automation: unofficial Chrome extension port of Jev Ultrafast where a Jev Choice picks the operation and DOM element each step and two Noul checks (goal reached, stuck) veto a premature DONE or BLOCKED, with a small text model used only when text must be typed.

  • Jevonian - Coding agents: local OpenAI / Anthropic / Responses-compatible proxy where one Jev call answers both the model route and the thinking level for jevonian/auto from session state, quota health, candidate capabilities, and cache-switch penalties; deterministic code filters candidates and owns every threshold first, a pinned model or explicit jevonian/<route> skips Jev entirely, and each decision lands in a local ledger with the serving model, reason, real token usage, and estimated cost.

  • jev-pruner - Context management: Claude Code plugin that trims long Bash output with Jev before the model ever sees it, keeping terminal noise out of the window.

  • jev-desktop - Computer use: supplies Jev action selection inside Codex Computer Use, choosing among desktop actions rather than asking a language model at every step.

  • jev-browser-bridge - Browser agents: plugs any CDP browser into a Jev loop, where a Jev Choice picks the operation and its target element each step from candidates read off the DOM rather than the layout, so the same agent runs on Chrome and on engines that never draw a page (Moli, Lightpanda, Kitesurf), passing at least 90% of runs on each of fourteen browsers tested.

  • Sedum - Browser end-to-end testing: in goal mode each turn asks one Jev Choice for the next operation (click, type, done or blocked) plus a speculative target among the page's offered elements, capped at 24 requests, 18 actions and 120 s, and the test passes only when an independent verify claim clears two Nouls (holds ≥ 0.75, contradicted flagged at ≥ 0.5), since the planner's done is never a verdict; authored-step tests reuse the same Choice to resolve each plain-English step, and with your own API key a 20-person team's PR suite costs $38–$91 a month vs $4,875 on a per-step AI platform.

  • killmyidea - Decision Tools: A startup-idea evaluation demo assigning KILL, FIX or SHIP labels from Jev scores.

  • jevify - Decision Tools: An Agent Skill for finding suitable Jev decision points and designing questions and comparison experiments.

  • hermes-jev - Decision Tools: An asynchronous Jev companion for Hermes Agent covering relevance, completion, recovery and optional admission decisions.

  • claude-jev - Decision Tools: Claude Code plugin: Jev for rule checks, verbatim compaction, and prompt routing

  • wechat-jev-assistant - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-chat-windows-deepseek-jev - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • Jev-chat-assistant - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-skill-router - Decision Tools: Claude Code plugin: asks TypeSafe Jev which installed skill fits each prompt and logs the answer (shadow-first). A working reference for the skill-suggestion cookbook on Claude Code — the README records why it is unlikely to help a strong model as a router.

  • jev-apply - Decision Tools: In Codex, Claude Code, or another CLI agent:

  • jev-bot - Decision Tools: Self-hosted Jev decision workbench and Feishu bot: automatic choices, probabilities, and experimental word/character writing.

  • dsh-jev - Decision Tools: DSH bundle that registers jev_ask for TypeSafe Jev noul, choice, and score answers.

  • jev-chat-windows-laya - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-demos - Decision Tools: Every demo lives in its own folder with its own README, dependencies and instructions. Most run without an API key in a clearly labelled `SIMULATED` mode; put `TYPESAFE_API_KEY=...` in the demo folder's `.env` for live results.

  • Jev-in-the-Loop - Decision Tools: Researching how Jev can accelerate tasks that rely on LLM decision-making.

  • jev-laya-benchmark - Decision Tools: Speed and accuracy benchmark: TypeSafe's Jev API vs the local Laya MLX typed-decision model on synthetic tasks

  • jev-predict-skill - Decision Tools: An Agent skill recipe that predicts another skill’s closed-set outcome from its rules and evidence.

  • astra-jev-harness - Decision Tools: The default `batch` policy retains an entire batch when any judgment is uncertain. The experimental `select --policy per-file` retains uncertain/unjudged files while omitting confidently irrelevant siblings. Dependencies and the global no-match fallback still apply. Compare before changing policy; fewer bytes alone do not establish correctness.

  • jev-codex-router-skill - Decision Tools: Portable Codex Skill for Jev model and reasoning-effort routing, with safe installation and Chinese usage guides

  • jev-playground - Decision Tools: A web playground for entering state and decision questions, then inspecting Jev answers and probability distributions.

  • fake-real-jev - Decision Tools: See link entry, the live scan timer, JEV's evidence-checking role, a saved REAL example, a saved FAKE example, and the linked sources. The live documentation scan shown ended without a verdict; its credit was returned. The coffee reports are clearly labeled saved examples, with their original analysis times.

  • jev_projects - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jev-crawlers - Decision Tools: Jev learns your repo's decision norms, then adversarially judges past decisions against them. Unix-style primitives (seed, expand, judge, verify, report, norms) with per-node typed judgments from typesafe-ai/jev.

  • jev-linter-action - Decision Tools: `glob` accepts one pattern or a newline-separated list. All matched files are reviewed together by default; `per-file: true` reviews each file independently. Missing inputs, malformed questions and ambiguous combinations fail before calls.

  • jev-no-enem - Decision Tools: Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.

  • jev-triage - Decision Tools: Millisecond-class test-failure triage for coding agents: RETRY / FIX_CODE / FIX_ENV, powered by TypeSafe Jev.

  • Jevatar - Decision Tools: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • jevchat - Decision Tools: A chat-style Jev demo whose answers are selected from predefined or custom options rather than generated prose.

  • JevCode - Decision Tools: JevCode - Jev can code. We want to dogfood JevCode

  • typesafe-jev-ruby - Decision Tools: Ruby client for Jev, TypeSafe's System One model: typed questions, probabilistic answers. Zero runtime dependencies.

  • xjevboost - Decision Tools: Add as much tabular data as you want to Jev models using adaptive ensembles that learn to query only the rows and columns needed.

  • turing-jail - Decision Tools: Interactive three-level AI interrogation game powered by TypeSafe Jev; write responses and pass plea, logic, and paradox verdicts to earn release.

  • Hermes JIT Context OS - Coding agents: uses Jev as a sub-millisecond System 1 Epistemic Gate and Domain Router to score AST relevance, test proofs, and tool targets, cutting autonomous agent turns by 31.3% and blind file exploration by 52.6% on SWE-bench with fail-open circuit-breaker resilience.

  • jev-agent-skill agent: Multi type: plugin - Developer tooling: Claude Code/ZCode skill that offloads classify/route, batch-screen, score, and compliance-check judgments to Jev via OpenCode Zen's free tier, bundling a zero-dependency jev.py caller (transient-500 retry, WAF-safe UA, GBK-pipe-safe stdin) and a production Taobao-shop comment-triage pipeline that keeps raw items out of the agent context.

  • Yappy - Computer use: macOS voice agent that asks Jev one Choice per step (operation and target control) over the front window's accessibility table, executes only validated high-confidence answers, and escalates to a full LLM agent on low confidence, no-effect actions, or unknown field values; author-reported 275–690 ms per decision.

  • JevLoop (parkavenue9639) - Agent runtimes: a Python runtime where Jev Choice decisions select tools and targets, uncertain decisions escalate to an LLM, and a shared guarded kernel supports isolated Docker workspaces and paired LLM-only comparisons.

  • GUI JEV Harness - Computer use: recursive screenshot grounding where Jev returns a Choice over grid-tile candidates at each level, and local probability and margin gates decide whether to descend or refuse, emitting only a raster point and bounding box and never clicking.

  • Visual-JEV - Multimodal models: Jev-style model built on Qwen3.5-4B that takes images directly, without first converting them to text.

  • DeepSearcher stopping-policy experiment - Agentic search: a standalone evaluation uses Jev Noul judgments on accumulated evidence to decide whether to stop or continue within a search-round budget, comparing stopping behavior, evidence recall, and decision cost.

  • jev-chat - Messaging: an Android accessibility service reads the conversation in WeChat, QQ, X, or Feishu, asks Jev Choice over candidate replies, and fills the draft box while sending stays manual; a Windows port does the same from offline OCR of the WeChat window.

  • SkillRanker agent: Claude Code type: cli - Coding agents: standalone Rust CLI that uses Jev to rank candidate skills against live session context, advising the next step through a Claude Code UserPromptSubmit hook.

  • AutoGPT - Autonomous agents: open-source autonomous agent platform featuring first-class TypeSafe Jev decision blocks for typed routing, filtering, scoring, and confidence-gated next-action dispatching.

  • oh-my-claudecode agent: Claude Code type: plugin - Coding agents: multi-agent team orchestration for Claude Code featuring opt-in Jev hooks for sub-millisecond judgment points, decision caching, and per-point egress controls.

  • jcode - Agent runtimes: RAM-efficient autonomous agent harness implemented in Rust with native TypeSafe Jev typed decision transport for memory pruning, browser navigation, and voice interaction routing.

  • opencode-jev-compaction - Replaces OpenCode compaction summaries with Jev keep/drop judgments that prune stale tool calls while preserving everything kept verbatim.

  • jev-auto-approve agent: Claude Code - Coding agents: Claude Code PreToolUse hook that asks Jev a Noul on whether a shell command is strictly read-only, auto-approving at 0.95 and otherwise falling back to the normal permission prompt without ever denying, while a local hard-no list and injection filter keep risky commands from reaching Jev; 0 of 8 state-changing commands were approved in its published calibration.

  • laya-browser-agent - Browser agent: derives each step from a Jev-shaped model — Laya through MLX or PyTorch, any duck-typed backend, or an arbitrary System One HTTP endpoint.

  • openclaw-jev-trigger agent: OpenClaw type: plugin - Agent automation: OpenClaw plugin and CLI that turn a plain-language --when / --not-when condition into a scheduled trigger script, asking Jev one Noul per tick through OpenClaw's decisionModel and waking the conversation model only when the condition becomes true at 0.7 or above; on 76 synthetic watcher ticks Jev was right on 75 with 0 false wake-ups at 231 ms p50 and about $0.000016 per check, against 87% for first-try JavaScript rules.

  • mu - Coding agents: a pi-based coding agent and desktop app that asks Jev Noul, Choice and Score questions at 38 decision points inside its loop, such as which chunks of a long tool output enter the context, whether a rule-flagged command was actually asked for, whether a fetched page or MCP result carries instructions aimed at the model, and whether a "done" was verified; each point acts on its verdicts by default, can be switched to shadow or off, and writes every verdict to a local ledger, and in the repository's replay benchmark goal-aware test-log selection by Jev cut 40-46% of the log without losing a required line.

Data Labeling & Curation

Source file: categories/data-labeling-curation.md

  • jev-curate - Dataset engineering: sifts synthetic JSONL and Parquet rows using Jev Noul checks and calibrated confidence scores, streaming passed records and rejections straight to disk.

  • typeful-triage - Open-source maintenance: multiplayer triage dashboard where Jev answers a fixed set of typed questions per issue — kind, severity, urgency, duplicate, and next step — and every human correction is kept and shown back to the model on later runs.

  • JevSpan - Information extraction: zero-shot named entity recognition that splits text at punctuation, asks Jev one Choice over every candidate window per entity type, verifies each nominee with a second Choice (the type, none, mixed or partial) and settles its boundary with a third, averaging 73.7 strict F1 across 12 Chinese and English NER benchmarks against 72.1 for direct extraction with Qwen3.8-27B.

  • GroundingJev - Visual annotation: a Jev-inspired Qwen3.5-0.8B model that maps an image and referring expression to four bounding-box coordinates in one forward pass, reporting an 8.61× inference speedup over its autoregressive base model.

Evaluation & Benchmarking

Source file: categories/evaluation-benchmarking.md

  • Jev Playground - Model evaluation: benchmarks Jev against Luna, Haiku, and Gemini at choosing validated legal moves in explicit-state games, scoring decision quality and consistency across a sequence of moves.

  • Jev vs Mistral and Gemini for event validation - Event discovery: head-to-head test of Jev against Mistral Small and Gemini Flash-Lite at validating local event listings.

  • jev-research-eval - Research automation: reproducible eval harness plus field note for Jev Ultrafast research-browser tasks, with QC'd cases, a suite runner, and a report generator.

  • Jev judge call vs dimension scores - Model evaluation: tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.

  • Jev Pong - Model comparison: Pong where the ball advances one step per model decision, putting Jev head-to-head with LLMs through Vercel AI Gateway.

  • Jev reranking is not a free win - Search reranking: a measured run over 33,047 catalog entries, 164 real queries, and 9,831 graded pairs reports that Jev reranking alone did not beat vector retrieval.

  • An early-access test of TypeSafe's Jev - Independent trial: measures calibrated judgments on early-access Jev and reports the resulting cost per decision.

  • jevcal - Model evaluation: fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, reports how much traffic still has to escalate to an LLM, and fails CI when a model update breaks the locked thresholds.

  • WindTunnel - Browser-agent benchmark: measures WebMCP against other browser-agent interfaces, with Jev appearing as one of the compared configurations.

  • jev-eval - Third-party check: compares Jev against GPT-4o-mini and Claude Sonnet 4.5 under identical conditions on the same judgment task.

  • minutes - Meeting notes: local-first transcription app whose live voice path runs its evaluations through Jev.

  • jev-orderby-bench - Model evaluation: measures whether a SQL ORDER BY over a Jev probability is defensible (pairwise inversion, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a pre-registered gate that jev-1.13.0 passes on 20 Newsgroups topics and fails four of six conditions on Amazon ESCI product relevance, and shows a DuckDB extension's default 40-row batching fails the ranking gate that one row per request passes.

  • jev-ood-calibration - Model evaluation: independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).

  • Odin R&D: Jev vs open-weight Laya - Model evaluation: runs the same 48 hand-written choice/score/noul rows through Jev's hosted /api/v1/systemone and a pinned open-weight Laya on a Mac (MLX, checked row-by-row against Laya's own reference code) under pre-registered refutation criteria, publishing the raw results record, 46/48 vs 39/48 accuracy, and a reproduce command.

  • latitude-llm - Evaluation & Observability: Latitude includes an optional Jev preclassifier for conversation checks and their selection records.

  • jev-review - Evaluation & Observability: A local MCP code-quality reviewer returning structured scores to coding Agents.

  • taskuary - Evaluation & Observability: An optional Jev judgment module in Taskuary for checking user-defined conditions on task state.

  • Canny - Evaluation & Observability: Keeps an execution ledger for Claude Code and Codex CLI to check for passing validation after edits.

  • jev-playground - Evaluation & Observability: A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.

  • typesafe-ai-benchmark - Evaluation & Observability: Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.

  • goodwatch-monorepo - Evaluation & Observability: A film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.

  • jev-benchmarks - Evaluation & Observability: A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.

  • jev-rerank-bench - Evaluation & Observability: Compares Jev, dedicated rerankers and chat models on the same retrieved passages.

  • jev-benchmark - Evaluation & Observability: Benchmarks Jev on chess moves and identifying which game NPC a player addresses.

  • jev-lm - Evaluation & Observability: A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.

  • jev-chat - Evaluation & Observability: A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.

  • jev-frontend-qa - Evaluation & Observability: Frontend QA that uses Jev to choose browser actions and checks contracts through DOM, HTTP and database evidence.

  • jev-behavior-study - Evaluation & Observability: An independent Jev 1.13.0 behavior study recording successes and failures across question framing, input conditions and games.

  • jev-exploration - Evaluation & Observability: A research repository tracking Jev claims and limitations, with calibration experiments and runnable examples.

  • jev-gomoku - Evaluation & Observability: A nine-board, 15×15 Gomoku workbench comparing how two Jev players respond to different input representations.

  • jev-agent-failure-benchmark - Evaluation & Observability: A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.

  • ask-jev - Evaluation & Observability: A Windows PowerShell tool for auditing recorded Codex execution evidence with :jev.

  • jev-synergy-screening - Evaluation & Observability: A Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.

  • hermes-jev-north-star - Evaluation & Observability: A Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.

  • jev-calibration-audit - Evaluation & Observability: Audits Jev calibration, option-wording effects and Korean judgments through public APIs and datasets.

  • jev-demos - Evaluation & Observability: Maze experiments comparing Jev single-step choices with multi-step lookahead.

  • jev-eval - Evaluation & Observability: Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.

  • jev-flash-review - Evaluation & Observability: An MCP review engine that evaluates Agent-supplied diffs against explicit rules.

  • foreman-jev - Evaluation & Observability: An experimental Jev supervisor for Codex workers with programmer-selected acceptance commands.

  • ASSAY-001 - Independent pre-registered check of Jev calibration and type safety on Banking77 / CLINC150, with a split verdict and full logs, written up at donttrustme.ai.

  • BTK audit studies - Content & growth: Jev striking-distance triage ranks SEO fixes and drives study pages; 1,204 pages judged per run, 4,816 judgments in under 3 minutes, $0.0048 per 12-query batch.

  • Can Jev Be a Better Agent Evaluator? - Agent evaluation: LangChain compares Jev against LLM judges on accuracy, repeatability, latency and cost, concluding Jev is the cheaper and more consistent judge for online evals.

  • jev-acento - Language evaluation: pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish state costs 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writing instructions in Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data.

  • Jevals.com - Model evaluation: independent leaderboard that asks Jev and six LLMs the same Noul, Choice and Score questions and grades every answer against human labels (PubMedQA, Banking77, HelpSteer2; 300 items × 5 runs each), finding Jev tied for first on PubMedQA yes/no at 1/28 of the top LLM's price, tied for second on Banking77 and no model beating the label base rates on HelpSteer2, with every per-decision probability published as CC BY 4.0 data.

  • Jev IDS - Network security: an Intrusion Detection System prototype that takes the metadata of a network flow and returns a verdict on whether it is an attack plus its threat category with probabilities, and on the NSL-KDD benchmark was 4.8× faster and 3.8× cheaper than a state-of-the-art LLM (GPT-5.6 Luna) while raising 15× fewer false alarms than a Random Forest model.

  • jev-test - Model benchmarking: reproducible test harness evaluating TypeSafe Jev Noul, Choice, and Score decisions via OpenRouter's Decisions API, comparing latency and accuracy against LLM prompt-and-parse baselines.

  • Jev vs Fable on 520 real social posts - Social media: a scheduler's pre-publish check asks Jev four Noul questions per caption (spam, clear opening, stands alone, promotional) as advisory signals, never a gate; on 100 posts labelled blind by Fable the two agreed 94/100 on promotion and 85/100 at a 0.65 spam threshold (Jev the stricter one 12 times to 3), and scoring all 520 posts cost $0.011 at a 341 ms median.

  • Jev Does Not Play Dice - Model evaluation: asks Jev a Choice over the six faces of a hidden fair die 400 times; Jev selects face 1 on all 400 trials with 82.9% mean reported probability and 19.0% accuracy, then tests whether stated probabilities survive in synthetic forecast documents, where a 30% shortage risk comes back as 5.3% via Choice and 26.7% via Noul; raw responses and analysis code on GitHub.

  • DecisionBench - Model evaluation: scores Jev Noul, Choice, and Score answers on pinned document-grounded tasks, counting malformed probability distributions as misses so model comparisons remain reproducible.

  • jev-regress-bench - Agent regression testing: after a config edit, one Choice (same / fact_differs / action_differs / specificity_differs) decides which of an agent's approved answers changed meaning rather than wording, and on 109 before/after pairs whose ground truth is derived from what each config rule does to the answer, Jev catches all 19 real changes with 13 false alarms against 33 for a markers-then-embeddings-then-LLM stack and 19 for the LLM judge alone.

  • jev-fanout-bench - Model billing: compares batched with one-question-per-call requests across 2,976 calls to jev-1.13-20260917 through OpenRouter's TypeSafe-compatible /systemone endpoint, reporting about 261 fixed input tokens per request, zero spread in the implied per-request cost across question counts, batched-vs-single answer differences comparable to repeat-request noise, and median input-token savings of 76–86% at eight questions.

  • SystemOneHarness type: cli - Model evaluation: execution harness and dual-loop test framework that compiles goals, browser environments, and MCP servers into bounded System One reflexes, evaluating Jev against deterministic baselines.

  • judgekit type: cli - Model evaluation: runs declarative YAML judgment tasks natively on Jev Choice/Score/Noul or any OpenAI-compatible backend (with a free rules fallback), gates low confidence at 0.7 (caught 3/3 misjudgments at 9% escalation, n=130), and publishes Chinese-scenario cost-accuracy numbers — 97.7% @ ¥0.105/1k decisions and 60.0% → 68.3% on a frozen 120-item human-labeled spam set at τ=0.10.

  • Convex Decision Evals - Model evaluation: asks Jev a Choice on 108 verified four-option questions about the Convex backend platform (no docs or tools in the prompt, each asked 3 times with shuffled options, random guessing 25%) alongside 14 LLMs, where jev-1.13 scores 84.6% at a 199 ms median and $0.0088 per full run against 98.0% at 2.12 s and $1.59 for the top model, with every answer, probability and raw request/response in a public explorer and the runner in get-convex/convex-evals.

  • jev-medhallu-benchmark - Medical AI: pre-registered test of Jev as a hallucination check on Stanford MedHELM's MedHallu (1,000 test items), asking one Noul on whether an answer misrepresents its PubMed abstract; Jev scored 92.9% against 92.4–95.1% for four fast LLMs at a 204 ms median and USD 0.03 per 1,000 checks, and letting Jev settle the 37% of items where it was at least 90% sure kept each LLM's accuracy with 37% fewer LLM calls.

  • zh-decision-bench - Benchmarking: first Chinese-language calibration benchmark for Jev-class decision models (378 items / 5 models incl. NeoHorse-Jev-4B; accuracy, ECE, option-order and zh-CN/zh-TW robustness; CC BY 4.0 dataset on Hugging Face).

  • jev vs. open alternatives - Document pipelines: compares Jev against open models and specialised tools on five chores — language detection, orientation, RVL-CDIP classification, bundle splitting and parse-tier triage — asked as Choice(2) up to Choice(16).

  • S1MB - Decision-model evaluation: compares Jev and open decision models across 137 English Choice, Noul, and Score benchmarks, including six synthetic generalization probes, with public evaluation data, recorded results, and an interactive leaderboard.

Calibration & Research

Source file: categories/calibration-research.md

  • decider - Open models: reproduces the System One shape with a Qwen3.5-2B fine-tune that emits typed decisions with calibrated probabilities in one pass.

  • openjev - Open research: independent local preview that answers bilingual probability questions from context, questions, and candidate answers, inspired by TypeSafe Jev.

  • Parallel Constrained Decoding (Qwen2.5-1B-RLCD) - Open research: RLCD-trained Qwen2.5-1B demo exploring open-source parallel constrained decoding as an alternative to Jev.

  • NanoJev - Open replica: a 0.6B parallel decision model that returns full probability distributions with no output-token decoding, shipped with its training pipeline, weights, and dataset.

  • open-alternative-jev - Open alternative: runs a Jev-shaped decision model locally on your own GPU.

  • mini-jev - Local reproduction: implements Jev's typed-decision interface on top of a local LLM.

  • Laya - Open alternative: non-autoregressive decision model that answers choice, score, and noul questions with RLCD-trained calibrated probabilities in a single ~35 ms forward pass, published on PyPI and Hugging Face.

  • Jev-compatible public API - Open research: a public Jev-shaped API backed by an open Qwen3.6-35B-A3B model so anyone can try the typed-decision interface.

  • kev - Trainable replica: a tiny Jev-like model on top of Qwen2.5-0.5B that trains and runs on a MacBook, shipped with its own research runs and evaluation scripts.

  • jevinci - Creative experiment: paints images by having Jev predict every pixel's colour in parallel, with predicted confidence deciding how wide each stroke is drawn.

  • jev-local - Local reproduction: Jev-compatible POST /v1/systemone server answering typed Choice/Score/Noul questions with confidence from open weights, verified as an official-SDK drop-in with temperature-fit calibration (set3 n=1316, 0.83 overall).

  • LitJev - Local reproduction: a reproduction of Jev that turns any Qwen model into a fast decision model, serving the same /v1/systemone schema (Choice, Score, Noul) with no training and no generated answer text.

  • ruling - Local reproduction: Jev-compatible POST /v1/systemone server that reads typed Choice/Score/Noul answers from the logits of any MLX checkpoint or OpenAI-compatible endpoint with no training, works as a drop-in for the official SDK, and replays Jev's published answers on 256 public judgments (231 vs Jev's 238, McNemar p = 0.21).

  • CUA-S1-FORMS - Specialist decision model: a 706,048-parameter, 2.8 MB jev-like option scorer that rates FILL / CHECK / CLICK / SKIP for each form field in one parallel pass, reporting 99.7% on its own form-filling eval against Jev's 83.6% - a specialist on home turf rather than a general win.

  • jevlike - Training library: build a small model that chooses among a changing list of text options and returns one probability per option in a single pass - the base CUA-S1-FORMS was built on.

  • jevbetter - Improved scorer: a stronger one-pass scorer over a variable list of text options, using a hashed n-gram encoder, rival-aware attention, and gated heads.

  • jevlike-esp32 - Edge deployment: exports a jevlike scorer as ESP32 firmware with a C scorer and a host-side check, putting one-pass decisions on a microcontroller.

  • von - Open alternative: a 395M non-autoregressive System One model that answers typed questions with calibrated probabilities in under 15 ms, positioned as a local drop-in replacement for Jev.

  • RSI-Jev - Trainable replica: an open 2B model answering Noul, Choice and Score on the same POST /v1/systemone wire format in one forward pass, over text or, since v4.0-VL, up to four images, researched and trained by a recursively self-improving AutoScientists loop that publishes every experiment it ran — 86.6 AUROC zero-shot on 2,162 VisA inspection photos (Jev-Omni 81.1, Gemma 4 12B 82.9) and 80.3% on five held-out image benchmarks.

  • Open Medical Jev - Medical evaluation: two frozen local readers answer one Noul-style yes/no probability per exam option, a fit-free router auto-releases items above the combined-confidence gate and escalates the rest, and a split-conformal candidate set bounds the error - landing within 2 points of hosted Jev on three 600-item national licensing exams with no fine-tuning, no distillation and no corpus.

  • OneJev - Open models: multimodal System One model in four sizes (0.8B to 27B); typed questions about a screenshot, photo, video or text get a calibrated probability for every option in one forward pass.

  • TetraJev - General decisions: two frozen open-weight readers give four readings per item (letter + per-candidate yes/no), fused fit-free and routed by agreement into auto-release or human review; evaluated across eight decision suites and a RAG reranking pass with published coverage–accuracy curves; no training, and it does not call the TypeSafe API.

  • Jebadiah - Open replica: Apache-2.0 decision models (27B, 9B, 4B on Qwen bases; bf16, GGUF and MLX) that answer Choice, Noul and Score questions with a probability for every option from one forward pass, and run anywhere: a standalone server with Jev's /v1/systemone wire and a playground, a llama.cpp script for the GGUF builds, or AINode (open-source local AI platform).

  • WebJev - Specialist decision model: Apache-2.0 open-weight Qwen3.5-35B-A3B fine-tune for browser agents, served by vLLM behind the same POST /v1/systemone and /api/alpha/decisions routes so a Jev client switches by changing only the base URL and key; inside the unchanged jev-ultrafast agent it completes 38.52% of 125 hand-picked real-website tasks graded by deterministic verifiers, against 16.67% for Jev 1.13.

  • JevForge - Open research: an end-to-end stack for auditable data construction, Qwen3.5-0.8B training, fixed Mind2Web and OOD evaluation, local serving, and a preliminary RLCD baseline.

  • Luce - Open recipe: describe the decision task in a sentence, an LLM teacher writes the training data, a LoRA + decision head on Qwen3-4B-Base answers choice/score/boolean questions with calibrated probabilities in one forward pass; trains on a 12 GB card, and reports accuracy and ECE next to Jev on identical test items (rule-generated tickets 91.1 vs 75.1, phishing 97.4 vs 62.6, GitHub issue priority 41.1 vs 37.5) with a browser replay demo that needs no GPU.

  • poorjev - Local reproduction: implements Jev's typed Choice/Score/Noul interface on commodity zero-shot NLI models and makes the confidence honest with temperature scaling and conformal abstention, shipping a reproducible calibration eval (ECE 0.170 to 0.071, cross-validated) that runs offline with no API key.

  • openJev-verdict-2.0 - Open decision engine: a calibrated 151M non-autoregressive model that reports beating both TypeSafe Jev and Laya on typed-decision benchmarks, shipped with its own test suite.

  • OpenDecision - Open alternative: a local semantic decision engine that describes itself as the open-source equivalent of Jev, answering Choice, Noul, and Score questions from structured state and documents without a hosted call.

  • TinyJev - Open alternative: a 596M pointer-head model that answers Choice, Score, and Noul in a single forward pass and returns calibrated confidence meant to be thresholded, so cases it is unsure about escalate to a human instead of being guessed; MLX-first on Apple Silicon, with a System One endpoint and weights on Hugging Face and ModelScope.

  • When a Judgment Layer’s Self-Reported Fields Lie - Independent measurement: tests Jev’s self-reported access-layer fields against ground truth rather than trusting them, reporting a verdict vocabulary reaching three values where the description lists six and a sufficient field that does not separate thin evidence from contradictory evidence; the contradiction reading has no JSON artifact behind it and the write-up says so in its own errata.

  • SemIf - Independent replication: reproduces Jev's typed-decision interface on open models, including an MLX backend on Apple silicon, and measures that typed decisions arrive together while a JSON answer streams token by token.

  • jev-verify - Developer tooling: recomputes Jev's confidence and expected-score identities against outputs published in public repositories rather than live API calls, separating vendor-channel examples (10/10) and recorded responses (843/854) from hand-authored fixtures (115/296), where all 121 outputs whose confidence equals the fractional part of their score are concentrated.

  • AnyJev - Open research: turns open LLMs into Jev-style decision models that read typed decisions and calibrated probabilities from next-token prefill distributions with zero fine-tuning, reducing order-flip rate and calibration error.

  • Verdict - Open alternative: Apache-2.0 118M multilingual bi-encoder that answers Choice, Score, and Noul questions on the same POST /v1/systemone wire format, calibrated with temperature scaling plus a split conformal abstain set with a coverage guarantee (ECE 0.01 to 0.03 on the public suites), runs on CPU or in the browser via ONNX, fits on your own labels in seconds, and its README says it loses to Laya on typed decisions (0.71 vs 0.77).

  • Jev-MedQA - Medical QA: a Jev-style implementation on Qwen3.5-4B that selects answers to text and image questions in one forward pass, reporting 69.42% accuracy versus 67.41% for standard generation with a 10.37x speedup across 153,889 questions from nine medical QA benchmark sets.

  • Jev Prime type: hosted - Text generation: a conversational agent with no language model, where every word is picked from ~4,700 options by Jev Choice questions one at a time, with confidence driving lookahead when the top pick falls below 0.65, beam search across sentence directions, and a self-critique loop that rewrites sentences scoring below threshold; live at talktojev.com, paper at doi.org/10.5281/zenodo.22940945.

  • CLM - Open alternative: an 8B System One model that answers the same Choice and Noul questions behind a TypeSafe-compatible API, matching Jev across computer-use, gaming and tool-calling with up to 9x lower latency and reporting 87.6% on Terminal-Bench 2.1 as a fine-tuned verifier.

  • Bespoke Nimble - Open alternative: a LoRA on Qwen3.5-9B that scores one allowed answer token per Choice, boolean or rubric-score question, released with its data pipeline, training config and eval harness under Apache-2.0, and reporting 90.1% on its 324-example holdout against Jev's 93.2%.

  • NeoHorse-Jev - Open alternative: Apache-2.0 4B decision model from TokenRhythm that answers Choice, Noul and Score questions via prefill-only inference on NeoHorse-1-4B, deployable with vLLM, SGLang or a native Python/CLI/HTTP runtime, scoring 77.70 across six text benchmark groups (highest among open-weight entries with complete results in its published comparison).

  • Jeff - Open alternative: MIT-licensed Qwen3.5 and Gemma 4 fine-tunes answering choice, noul and score on the same /v1/systemone format at about 22 ms per decision, published with a panel that measures Jev itself at 0.828 accuracy and 0.053 ECE while stating it claims no statistical significance.

  • AutoJev - Open recipe: a 27B multimodal decision model trained with full-weight SFT on 73,000 examples over one H200, serving choice, noul and score on /v1/systemone with per-checkpoint provenance and calibration plots released.

  • Lev - Open alternative: a 4B LoRA on a Qwen backbone published as a System One decision model and tagged for calibrated decisions, classification, routing and moderation.

  • openjev - Open implementation: ranks a Doom action menu with one /score call per step and supports per-task fine-tuning, released under MIT with the terminal run recorded.

  • Jev-Omni - Multimodal System One: a Gemma-4-12B fine-tune that answers typed questions over text, images, audio and video, at 1,402 downloads in its first ten days.

  • Bekko System One - Open decision models: independent 17M–400M English models for Choice, Noul, and Score, with public weights, training code, and an ONNX browser demo; v0 remains substantially behind Jev on the project's generalization tests.

  • jevcrypto type: library - Creative experiment: a crypto.randomUUID() look-alike npm package that writes Jev's raw Noul probabilities on 15 code-point permutations of any prompt, plus one Choice for the variant digit, into the bytes of a UUID v4-shaped string; deliberately not cryptographically random.

  • Vev - Open alternative: an open-source Jev implementation with vision input, LoRA fine-tuned on Qwen3.5-4B and 9B, serving Choice, Score and Noul questions on the /v1/systemone wire format with screenshots and photos placed directly in the state so one decision can use both text and image; weights are CC BY-NC 4.0, non-commercial only.

  • WaterSheep - Open alternative: an Apache-2.0 open-weight model that answers noul, choice, score and multi-label questions with a calibrated probability for every option, serves POST /v1/systemone so TypeSafe's Python SDK works after a base-URL change, and has an ONNX build that runs in the browser.

Infra / SDKs / Integrations

Source file: categories/infra-sdks-integrations.md

  • eve - Agent frameworks: Vercel's eve engine ships Jev as the default evaluation model (typesafe-ai/jev) in its experimental evaluate path.

  • AI CLI - Developer tooling: Vercel Labs CLI that can run Jev as the evaluation model for its evaluate command.

  • jev-mcp (jkudish) - MCP ecosystem: proof-of-concept MCP server that puts Jev claim verification, content screening, and candidate ranking behind standard MCP tools.

  • jev-mcp (blakestone-x) - MCP ecosystem: MCP server exposing Jev classify, score, check, match, and screen as tools for any agent, with confidence on every answer.

  • zio-typesafe-ai - Scala ecosystem: ZIO client for TypeSafe AI with a typed DSL over Jev decisions.

  • TypeSafe AI Swift SDK - Swift ecosystem: dependency-free Swift 6 client for Jev Choice, Score, and Noul questions with strict concurrency, configurable authentication and retries, and offline transport tests.

  • laravel-typesafe-jev - PHP ecosystem: unofficial Laravel integration for Jev with typed responses, async requests, scoped dependency injection, and testing fakes.

  • advocaat - Data tooling: small type-safe client for asking Jev questions about a dataset.

  • jevclient - Python ecosystem: async client for Jev published on PyPI.

  • LlamaIndex Jev - Retrieval / RAG: unofficial LlamaIndex adapter where Jev Scores each retrieved passage and Choice/Noul selects the query engine, with nfcorpus nDCG@5 0.340→0.396 at about $0.0003/query.

  • safer-with-jev - Cloud infrastructure: Neon Function proxy for the Neon AI Gateway that routes decisions with Jev.

  • typesafe-ai/skills - Official tooling: installable agent skills package (npx skills add typesafe-ai/skills) that teaches agents the Jev workflow.

  • Smithers - Agent frameworks: TypeScript workflow framework with a Jev session checker wired into its workflows.

  • skillbox - Skills infrastructure: self-hosted versioned skills library that adds optional Jev recommendations using your own TypeSafe or Gateway key.

  • Jevbridge - Agent bridges: ACP and MCP adapter that exposes Jev typed decisions to Codex, Claude, Grok, and other LLMs.

  • jev (Elixir) - Elixir ecosystem: GenServer client that replies with Jev's answer so callers can pattern match on it directly.

  • jev-go - Go ecosystem: community Go SDK for Jev.

  • jev-cli - Developer tooling: small dependency-free CLI for Jev.

  • decide-mcp - MCP ecosystem: configurable decision server with percentage scores and bias-profile routing on top of Jev.

  • typesafe-jev-examples - Starter examples: worked ticket-triage and reranking examples runnable through OpenRouter without an early-access key, shipped with their own sample data and Makefile.

  • ai-python - Python ecosystem: the official Vercel AI SDK for Python carries Jev through its evaluation operation and Gateway examples.

  • Cline plugins - Coding agents: Cline's official plugin collection includes a Jev-driven browser plugin (jev-browser), so Jev arrives as a first-class Cline capability.

  • hono-jev-router - Web frameworks: Hono middleware that routes HTTP requests by meaning rather than by method and path, deciding with Jev.

  • rotom - Local gateways: OpenAI- and Anthropic-compatible API gateway that carries Jev through its model catalog and evaluation path.

  • Jev AI - Developer tooling: public Jev playground and API that puts typed Choice, Score and Yes/No questions to the model about pasted text - ticket triage, moderation, review scoring - and returns a parsed answer with a confidence value in about 0.5 s per decision.

  • jevql - Data tooling: psql-shaped CLI and Go/TypeScript/Python SDKs that run plain SQL on a vanilla Postgres (no extension) and then ask Jev Noul, Choice, or Score questions about each surviving row so the client can apply jev() filters, jev_prob sorts, and jev_choice groups.

  • sqlite-jev - SQLite ecosystem: loadable C extension and Python package that expose Jev Noul, Choice, and Score judgments as SQL functions and batched virtual-table queries with confidence results.

  • jevkit - Developer tooling: Rust CLI that validates Choice/Score/Noul question sets with 13 offline lint rules before any Jev call, then sends the canonical wire payload and prints parsed, confidence-bearing JSON answers to stdout using exit code 2 to reject a billed-but-useless request.

  • jev-use - MCP ecosystem: Claude Code / Codex / pi plugin (MCP server + library, native pi extension) that hands agent steps needing no text output to Jev as typed judgments — untypeable and generation-needing questions are rejected before the call, low-confidence answers come back flagged as priors, and a fail-open PreToolUse gate can only deny or ask.

  • huncho - TypeScript ecosystem: dependency-free SDK that turns Jev Noul, Choice and Score answers into named decisions with enter/exit thresholds (hysteresis), nested decision trees settled in one call, a JSONL journal, replay of a threshold change over recorded answers with no inference, and Brier/reliability calibration, over TypeSafe direct, OpenRouter or Vercel AI Gateway.

  • jev-experiments - Demo collection: 22 latency-focused Jev applications built by Devin, each with its own README and testing notes, spanning shell guards, log sentinels, instant search, reranking, and voice turn-taking.

  • ruby_decision_model - Ruby ecosystem: client for decision models such as Jev, so Ruby applications can put typed questions directly to the model.

  • s1_ruby - Ruby ecosystem: makes System One measurement, and the collapse that follows it, a Ruby primitive, with a TypeSafe provider behind its own spec suite.

  • kojev - Kotlin ecosystem: Kotlin Multiplatform (JVM, Android, iOS) client for Jev that answers Choice and Score questions as the caller's own enums, with one typed way to read answers, no default thresholds, and offline MockEngine tests.

  • hunch - Python and TypeScript ecosystem: libraries that turn Jev Choice, Score, and Noul questions into functions over lists and DataFrames (classify, score, check, where, extract, pick, rank, verify), with request deduplication, caching, and optional escalation of unsure rows to an LLM that must pick from the same labels; TypeScript port at hunch-js.

  • stuntd - Local runtime / learning proxy: Jev-compatible local server on Laya that also proxies a Jev upstream, records every Choice, Score and Noul decision, trains a head per decision site, and answers live with calibrated confidence, falling back to the upstream below its threshold.

  • should-i-jev - Migration tooling: dependency-free CLI that scans LLM logs and code for decision-shaped calls, prices the Jev migration, asks a Jev endpoint which call sites to take (--jev-selfcheck), calibrates typed answers against ground truth (ECE, reliability, risk-coverage), and generates a reviewable migration PR with Choice/Score/Noul map sketches.

  • vellum-assistant - MCP & Integrations: An optional Jev provider in Vellum Assistant sends conversation state and explicit questions to TypeSafe.

  • typesafe-mcp - MCP & Integrations: An MCP server that lets Claude Code, Claude Desktop, Codex and Pi ask Jev typed questions.

  • pi-typesafe - MCP & Integrations: A Pi Jev extension providing a decision tool, terminal playground commands and an API for other extensions.

  • Jevbridge - MCP & Integrations: Exposes a shared structured-decision interface for Jev and other models through ACP, MCP and a CLI.

  • synkora-ai - MCP & Integrations: Synkora includes optional TypeSafe client tools for classification, scoring and yes/no judgments.

  • jev-judge-mcp - MCP & Integrations: Typed judgment tools for MCP agents. TypeSafe's Jev model as verify, screen, find, classify, rerank, decide, compare, extract, review, gate, and score: the model judges, policy decides auto, review, or escalate.

  • plasmallm - MCP & Integrations: A Jev Decisions adapter in a KDE Plasma assistant widget for displaying structured judgments.

  • jevwire - MCP & Integrations: Provides Jev MCP tools, an embeddable library and Claude Code hooks for Agents.

  • harness-router - MCP & Integrations: Fast decision routing for agent harnesses — native MCP with Jev for tool selection and MCTS for multi-step decisions.

  • jev-codex-plugin - MCP & Integrations: Open-source Codex plugin for TypeSafe Jev decision consultation, failure diagnosis, and evidence-based completion review

  • jev-mcp - MCP & Integrations: An MCP server and Claude Code plugin for Jev classification, scoring, checks and batched questions.

  • jev-classifier - MCP & Integrations: A local Jev tool-routing gateway for coding Agents, with an MCP suggestion interface.

  • tenbin - MCP & Integrations: Documentation, an MCP server and a Skill for designing Jev judgments with question linting, batch evaluation and calibration.

  • jev-agent-kit - MCP & Integrations: jevkit: fast typed decisions for agents. CLI and MCP tools (route, triage, guard, grep, rank, compact, judge) on TypeSafe Jev. Zero dependencies.

  • Jev-AI-Skill - MCP & Integrations: One AI skill + MCP server for Claude Code, Codex and Hermes: Jev (TypeSafe) gates large-model turns (event triage, review verdicts, owner questions, tool choice), runs build-and-repair loops with independent review, routes models and skills, and guards Git steps. Script-only watch mode and a shared HTTP server.

  • jev-as-quant - MCP & Integrations: Typed System-1 decisions (Laya/Jev) as the judgment layer of a quant research stack, with Claude as System 2. Requirements → design → code → experiments.

  • jev-mcp - MCP & Integrations: An evaluation-focused Jev MCP server for individual questions, batch processing and question or threshold comparisons.

  • jev-mcp-server - MCP & Integrations: MCP server for Jev (TypeSafe System One): the three official question types — choice, score, noul — plus batch classify. Calibrated probabilities, ~0.5s, <$0.001/call.

  • jev-skill-router - MCP & Integrations: Keep skill catalogs outside the main LLM context. Jev selects relevant skills through one read-only MCP tool.

  • jev-workbench - MCP & Integrations: Defines, tests, and publishes Jev decision functions in a local UI so backends and Agents can call fixed versions.

  • jev-agent-toolkit - MCP & Integrations: Jev-first portable Agent Skill and optional MCP bridge for Claude Code, Codex, Cursor and compatible agents.

  • jev-mcp - MCP & Integrations: MCP server wrapping TypeSafe's Jev System One models — typed noul/choice/score judgments for AI agents

  • jev-mcp - MCP & Integrations: Local MCP server exposing TypeSafe Jev (System One decision model) as native Claude Code / Codex tools

  • jev-mcp - MCP & Integrations: MCP server for Jev (TypeSafe AI's System One model) — give any agent typed, calibrated decisions: classify, score, check, gate risky tool calls. Try free: jevtypesafeai.com

  • jev-mcp-spring - MCP & Integrations: On success each tool returns its typed result directly. On failure the tool call fails at the MCP protocol level (`isError: true`) with a short, safe message — no response body or stack trace is ever echoed back.

  • jev-playwright-mcp - MCP & Integrations: Jev-augmented Playwright MCP proxy — page-state triage, prompt-injection shielding, goal-based snapshot pruning, risky-action gating. Drop-in wrapper around @playwright/mcp for any coding agent.

  • jev-rust-review - MCP & Integrations: Rust-aware code review for Claude Code and coding agents, powered by TypeSafe Jev

  • jev-toolkit - MCP & Integrations: MCP-first toolkit for TypeSafe/Jev — the System One decision model. One stdio server (jev mcp) serves any MCP-capable harness, backed by one local event log and Prometheus impact metrics you can scrape into your own Grafana.

  • n8n-nodes-jev - MCP & Integrations: Jev by TypeSafe AI for n8n: typed decisions, probabilities, and confidence-aware workflows

  • n8n-nodes-typesafe-jev - MCP & Integrations: An n8n community node for submitting typed questions to TypeSafe Jev.

  • toolJev - MCP & Integrations: Code Mode for MCP, where the sub-model is a calibrated decision model (Jev), not an LLM. Benchmarked on MCPToolBench++, LiveMCPBench, When2Call and live Claude agents.

  • typesafe-jev-mcp - MCP & Integrations: This repository is an MCP server for TypeSafe Jev that provides a single evaluate tool taking state plus typed questions and returning noul, choice, or score answers with probabilities.

  • typesafe-jev-opencode - MCP & Integrations: Jev is not a conversational replacement for Gemini, Claude, or GPT. It evaluates application state against typed questions and returns structured answers and probabilities that an agent can use to route or gate work.

  • jev_ampcode - MCP & Integrations: An Amp plugin for comparing supplied alternatives against supplied evidence and priorities.

  • jev-eyes - MCP & Integrations: Also in the state: `image` (size, source), `blocks` (`[x, y, w, h]` boxes with OCR confidence) and, if installed, `labels`. `see(img, compact=True)` keeps only `image`, `text` and top label names when tokens matter more than positions.

  • jev-in-mcp - MCP & Integrations: MCP relay that adds use_jev to every server: Jev picks the tool calls, the calling model writes the values Jev cannot choose, the relay executes. Built on jev-dev-kit.

  • jev-mcp - MCP & Integrations: MCP local que expone Jev (TypeSafe) como herramienta para Claude Code, Codex, Hermes y cualquier agente: ask_jev y list_jev_models, sin dependencias

  • jev-mcp - MCP & Integrations: A Rust MCP server for TypeSafe AI Jev structured decisions

  • jev-mcp - MCP & Integrations: `jev-mcp` exposes TypeSafe AI's Jev decision model as four conservative, read-only MCP tools for **bounded probabilistic decisions**.

  • jev-mcp-open-source - MCP & Integrations: Self-hosted Jev MCP on Cloudflare Workers with intent routing, retrieval reranking and batch judgments

  • jev-review-mcp - MCP & Integrations: Single-purpose MCP server (one tool, one job): a code-review gate powered by TypeSafe Jev (System One decision model).

  • jev-routing - MCP & Integrations: Go Jev harness for Claude Code, Codex, and Grok Build. No npx. Not an MCP server.

  • jev-screen-mcp - MCP & Integrations: Single-purpose MCP server (one tool, one job): a content-moderation gate powered by TypeSafe Jev (System One decision model).

  • jev-tool-search - MCP & Integrations: Tool search for LLM agents: BM25 vs embeddings vs rerankers vs Jev on 525 real MCP tools, plus an experimental Jev search engine

  • jevmod - MCP & Integrations: Moderation for communities and apps: every message gets a probability for **spam, scam, harassment, nsfw, off-topic, self-harm, doxxing, sexual content involving minors**, and for **rules you write in plain English**. You set the thresholds and the actions. Every decision is logged with its numbers.

  • mcp-server-jev - MCP & Integrations: Typed AI decisions for Codex, Claude and any MCP client, powered by TypeSafe Jev. Classify, score and evaluate with one generic tool.

  • openclaw-typesafe-ai - MCP & Integrations: An independent OpenClaw plugin registering one explicitly invoked typesafe_decide tool.

  • openrouter-jev-mcp - MCP & Integrations: A Python decision gateway and Model Context Protocol (MCP) server exposing TypeSafe's Jev model through OpenRouter's decisions endpoint to Claude Code, Codex, and Cursor agents.

  • composio - SDK & Decision Frameworks: An optional TypeSafe provider for Composio that uses Jev to choose among tools and bounded argument options.

  • ai - SDK & Decision Frameworks: The TypeSafe provider in AI SDK lets TypeScript applications call Jev through the shared evaluate interface.

  • eliza - SDK & Decision Frameworks: An optional TypeSafe HTTP adapter in Eliza’s source, not registered with the Agent runtime by default.

  • langchainjs - SDK & Decision Frameworks: An optional LangChain.js TypeSafeClassifier integration for sending state and predefined questions to Jev.

  • rig-typesafeai - SDK & Decision Frameworks: An experimental TypeSafe crate in Rig for expressing Jev questions and answers with Rust types.

  • jev - SDK & Decision Frameworks: jevos is an open-source, Jev-compatible alternative to TypeSafe's Jev for yes/no decisions that runs entirely on a laptop CPU. Send text plus a yes/no question over HTTP and get back P(yes) from a single forward pass of a 1B model — no text generation, no GPU required.

  • req_llm - SDK & Decision Frameworks: A TypeSafe provider for calling Jev through ReqLLM’s evaluate interface in Elixir.

  • simple-jev - SDK & Decision Frameworks: Adapter turning open LLM endpoints into Jev-compatible classification services without training a separate classifier head.

  • jev-skill - SDK & Decision Frameworks: This project is a collection of Jev use cases, workflows, and agent skills with a stdlib-only Python decision wrapper.

  • openjev - SDK & Decision Frameworks: An independent System One decision server compatible with Jev’s API, running an open DiffusionGemma model.

  • instructor-php - SDK & Decision Frameworks: A TypeSafe Decision driver within Instructor PHP’s Polyglot module.

  • jev-visual - SDK & Decision Frameworks: Local Jev-like visual inference experiment on Apple Silicon Mac. Scores and classifies single images across multiple questions with 3 playable game demos.

  • jev-dsh-decision - SDK & Decision Frameworks: Provides Jev structured decision support for Agent Harness to recommend tools, Skills and Agents and return judgments with probabilities, with a native DeepSeek Harness plugin and an iPolloWork entry serving OpenCode, DeepSeek Harness and Codex Harness.

  • pi-fabric - SDK & Decision Frameworks: Pi’s programmable runtime includes an optional Jev loop for observing state, making decisions and running bounded actions.

  • typesafe-sdk-js - SDK & Decision Frameworks: The JavaScript and TypeScript SDK published by TypeSafe, with typed Jev requests and answers.

  • openai-scala-client - SDK & Decision Frameworks: A dedicated TypeSafe module in a Scala client that supports multiple AI providers.

  • typesafe-sdk-python - SDK & Decision Frameworks: Official TypeSafe Python SDK with synchronous and asynchronous clients for Jev System One, plus question and answer types.

  • jevbench - SDK & Decision Frameworks: JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open.

  • runline - SDK & Decision Frameworks: A TypeSafe plugin exposing Jev decisions as callable actions in Runline Agent JavaScript.

  • ai - SDK & Decision Frameworks: A Jev forwarding endpoint in the Hack Club AI proxy, using its authentication, limits and usage logging.

  • effect-agent - SDK & Decision Frameworks: An Effect Agent TypeSafe decision provider for typed question sets and optional model selection.

  • jeview - SDK & Decision Frameworks: An unofficial local visualizer for Jev (TypeSafe): a live view of every call your code makes. Not affiliated with TypeSafe AI.

  • Jev - SDK & Decision Frameworks: Unofficial TypeSafe Jev showcase — System One decisions, not chat.

  • solar-mini4-jev - SDK & Decision Frameworks: A drop-in wrapper that exposes Upstage **Solar Mini4** through the TypeSafe Jev System One API shape.

  • ask-jev-skill - SDK & Decision Frameworks: A skill for Hermes and other agents to query TypeSafe Jev for bounded option judgments and confidence escalations.

  • go-jev - SDK & Decision Frameworks: Go SDK and CLI for TypeSafe Jev: typed decisions (yes/no, choice, score) from a model

  • jev-capability-atlas - SDK & Decision Frameworks: This repository collects real Jev API-call receipts, test suites, and bilingual guides to map which narrow-decision tasks suit Jev and how Agents should evaluate and report fit.

  • minojev - SDK & Decision Frameworks: Decisions, not tokens: minojev reads calibrated, typed probability distributions straight from hidden states in one forward pass — zero output tokens, fully reproducible on a laptop CPU.

  • jev-agent-design-with-topk-logits-choices - SDK & Decision Frameworks: Research design for a Jev-native agent system: tool integration, speculative parameter proposals, external helper logits Top-k proposals with Jev-controlled fallback ,decision-aware hierarchical memory, and dependency-aware replanning.Feature:Jev naturallanguage conversation prototype using external helper logits and dynamic Top-k token selection.

  • jevalyn - SDK & Decision Frameworks: The decision layer for your Rails app. A Rails-native wrapper around TypeSafe's Jev System One API: typed, calibrated decisions in your control flow.

  • jev-to-answer - SDK & Decision Frameworks: This project integrates Jev to provide structured decisions for its workflow. See the repository for implementation details.

  • swift-jev - SDK & Decision Frameworks: A Swift client for TypeSafe AI's Jev — typed judgements, not text

  • swift-typesafe - SDK & Decision Frameworks: A community Swift TypeSafe client with typed questions, dynamic questions and response parsing.

  • discern - SDK & Decision Frameworks: A semantic control-flow library for Effect: Jev's `Choice` / `Noul` / `Score` answers become typed branches. Thresholds are caller-supplied, and an answer below them takes an explicit `Uncertain` branch the compiler forces you to handle rather than being rounded up to the top label. Procedure routing filters candidates with deterministic predicates first and skips the model call entirely when one candidate survives.

  • jev-rs - SDK & Decision Frameworks: System One judgments (noul/choice/score) from any LLM in one prefill — a Rust, Jev-compatible /v1/systemone engine

  • typesafe-ai - SDK & Decision Frameworks: A Rust TypeSafe client with asynchronous reqwest or blocking ureq backends and observable retries.

  • learn-jev-end-to-end - SDK & Decision Frameworks: **Learn Jev end to end** is a free, hands-on course. In 12 short notebooks you go from *"what is Jev?"* to building **13 real AI tools** with it: an email triage job, a scam-text detector, a code vulnerability hunter, an agent safety guard and more. You need **one API key**, and running the whole course costs **less than $0.20**.

  • typesafeai-dotnet-sdk - SDK & Decision Frameworks: Community .NET SDK for TypeSafe AI and Jev, providing asynchronous typed evaluation for Choice, Score, and Noul primitives.

  • jev-dspy-lab - SDK & Decision Frameworks: Reproducible calibration and selective-risk benchmarks for Jev/TypeSafe decisions in DSPy workflows

  • typesafe-sdk-java - SDK & Decision Frameworks: A community Java TypeSafe client with a Spring Boot Starter for configuring Jev calls.

  • jev-ai-sdk-form-router - SDK & Decision Frameworks: Route form submissions to the right people with Jev and AI SDK.

  • SpecPi - SDK & Decision Frameworks: A Pi configuration and extension bundle with an optional Jev advisor for capabilities and workflow checks.

  • typesafe-sdk-go - SDK & Decision Frameworks: A Go TypeSafe SDK for defining typed questions and reading Jev choices, scores, and probabilities.

  • zod-jev - SDK & Decision Frameworks: Adds semantic rules to Zod validation, such as checking whether text matches a description or contains personal information.

  • hermes-jev-plugin - SDK & Decision Frameworks: TypeSafe Jev (System One) decision tools for Hermes Agent: jev_check / jev_route / jev_score / jev_evaluate

  • jev-dsl - SDK & Decision Frameworks: An early Haskell DSL that describes labeled Jev questions, renders requests and decodes matching answers.

  • jevgo - SDK & Decision Frameworks: Unofficial Go client for the TypeSafe AI System One API (Jev) — typed questions in, calibrated answers out.

  • JevOps - SDK & Decision Frameworks: Jev is a **gate**, not a generator. This package does **not** write Lean. Lake (or another oracle) lives in the implementation that *uses* the kernel.

  • typesafe-sdk - SDK & Decision Frameworks: A community Ruby client for TypeSafe System One, defaulting to jev-latest.

  • daf-jev - SDK & Decision Frameworks: A Python toolkit for Jev requests, batch evaluation, calibration and MCP access.

  • jev_jsonschema - SDK & Decision Frameworks: `probabilities` is keyed by your schema's values, not Jev's internal labels, so a score of `1`–`5` reads as `"1"`–`"5"` and not `"0"`–`"4"`. Noul questions carry no confidence of their own, so `confidence` is `None` for booleans and numbers.

  • open-bonsai-jev - SDK & Decision Frameworks: openjev's mechanism, Bonsai's weights: typed decisions read straight from one forward pass of a 1.75-bit 27B model. Credit to TheoLeeCJ (SemIf/OpenJev) and PrismML.

  • typesafe-ai-rs - SDK & Decision Frameworks: An independently maintained Rust SDK with async and blocking clients, retries and response metadata.

  • jev-android - SDK & Decision Frameworks: A Kotlin Android SDK for UI automation powered by TypeSafe Jev, with an accessibility runtime and sample app.

  • jev-go - SDK & Decision Frameworks: A Go TypeSafe System One client with typed questions, answers and batching helpers.

  • jev-sdk-java - SDK & Decision Frameworks: Type-safe Java 21 client for the TypeSafe AI Jev (System One) decision API

  • jevriel - SDK & Decision Frameworks: **Codex · Claude Code · Portable skill** | [Apache-2.0 code and docs](LICENSE) | Early release

  • typesafe_sdk - SDK & Decision Frameworks: An Elixir TypeSafe SDK that brings Jev’s typed questions and probabilistic answers into Elixir applications.

  • usejev - SDK & Decision Frameworks: Run Laya locally with Bun: native ONNX inference, a TypeSafe-compatible API, and a bilingual decision playground.

  • goodall - SDK & Decision Frameworks: An optional TypeSafe package in a Go Agent library, using Jev as a tool or routing judge alongside chat models.

  • jev - SDK & Decision Frameworks: A Go client for the TypeSafe AI's System One API and its model, Jev.

  • jev-java - SDK & Decision Frameworks: Unofficial Java SDK for TypeSafe Jev and Vercel AI Gateway, with Spring Boot and WebClient support

  • jev-pilot - SDK & Decision Frameworks: Fast System-1 Decision, Arbitration & Safety Engine for Autonomous AI Agents (Powered by TypeSafe Jev)

  • jevlang - SDK & Decision Frameworks: The simplest way to write decision workflows in Python. Python with a smart if.

  • questions - SDK & Decision Frameworks: A TypeScript decision library that asks typed questions via Zod or native batches, defaulting to TypeSafe Jev, with optional Vercel or generative adapters.

  • ask-jev-ai - SDK & Decision Frameworks: Most AI demos generate text. Jev does not. It reads a sentence and returns typed answers with probabilities: a choice, a yes or no, a score. That makes it usable as a primitive inside ordinary code rather than a chatbot bolted onto a page.

  • jev-does-not-play-dice - SDK & Decision Frameworks: Code, recorded outputs, and analysis scripts for probability-output experiments with Jev: fair random draws, Noul (Yes/No) questions, and forecast documents.

  • jev-layer - SDK & Decision Frameworks: Portable System-1 decision layer for agent harnesses with host-owned routing, receipts, replay, and fail-open integrations.

  • jev-numeric - SDK & Decision Frameworks: **Both are multiway decision trees; decimal-digit decoding is a ten-way instance.** On an aligned decimal grid, they can have identical branches and leaves, expressed through different prompts. The digit is a **Choice option**, not a vocabulary token; Jev returns the option probabilities directly.

  • Jev4Mellea - SDK & Decision Frameworks: This is an unofficial, synchronous adapter. Jev evaluates text; it does not generate or repair it.

  • jevex - SDK & Decision Frameworks: An Agent experiment where Jev directs a tool loop, a chat model fills arguments and prose, and MCP tools execute.

  • typesafe-ai-rails - SDK & Decision Frameworks: Ruby on Rails integration gem for TypeSafe AI and Jev, providing model-level classification and decision policy patterns.

  • typesafe-sdk-rust - SDK & Decision Frameworks: A Rust client for TypeSafe with asynchronous and optional blocking calls plus typed question and answer wrappers.

  • dsh-jev-decide - SDK & Decision Frameworks: This DSH plugin registers TypeSafe Jev as an Agent tool that returns calibrated probabilities for noul, choice, and score judgments without generating text.

  • everything-about-jev - SDK & Decision Frameworks: tell you everything about jev,TypeSafe AI's System One model for typed decisions.

  • jev - SDK & Decision Frameworks: Ruby client for the typesafe AI Jev model

  • jev-benchmark - SDK & Decision Frameworks: Rubric-Based Zero-Shot Classification Benchmark: Jev vs Claude Haiku 4.5 vs Claude Sonnet 5 vs OpenJev on rubric-conditioned classification, chained decision execution, and exam grading -- with full price tracking.

  • jev-builder - SDK & Decision Frameworks: A browser form for building requests to TypeSafe's Jev: pick a template, fill in the blanks, copy the request. No JSON, no install, runs locally.

  • Jev-Persian-Benchmark - SDK & Decision Frameworks: Benchmarks for **Jev** on **480 authored general Persian questions** and a **24-excerpt classical Persian poetry pilot** (48 main questions plus 48 controls). Related general questions are batched; poetry questions run individually. Raw responses are saved and answers are scored locally, without a runtime model judge. The general benchmark's historical Laya comparison is retained below.

  • Jev-PhoneControl - SDK & Decision Frameworks: Visual Android automation powered by three agents: vision, a text-only supervisor, and TypeSafe JEV. Executes actions through ADB with a local web console.

  • jev-skills - SDK & Decision Frameworks: Practical agent skills and examples for building with Jev. API setup, routing, ranking, and evidence checks.

  • jev-web-analyzer - SDK & Decision Frameworks: See what Jev thinks about your SaaS website — powered by ReplyNodes web context and Vercel AI Gateway.

  • jevclient - SDK & Decision Frameworks: An asynchronous Python Jev client that batches typed questions in one request.

  • jevgo - SDK & Decision Frameworks: A community Go client with a standard-library core and optional Langfuse tracing.

  • jevrag - SDK & Decision Frameworks: Replaces hardcoded RAG thresholds with explicit calibrated decision points. Five primitives (retrieval stopping, chunk splitting, context selection, answer abstention, cache trust) behind one swappable state → Decision → confidence → action interface, each evaluated on real datasets with a calibration harness that reports honestly.

  • qualm - SDK & Decision Frameworks: A TypeScript wrapper for Jev decisions with an explicit unsure branch.

  • typesafe-go - SDK & Decision Frameworks: An unofficial Go client without third-party dependencies for System One calls and model discovery.

  • typesafe-sdk-php - SDK & Decision Frameworks: A community TypeSafe SDK for PHP 8.2+, with synchronous calls and Guzzle-based async requests.

  • jev - SDK & Decision Frameworks: Unofficial Go client for TypeSafe's System One API and its model, Jev.

  • jev_dart - SDK & Decision Frameworks: Jev Dart SDK to build cli, server and Flutter apps.

  • jev-architecture-research - SDK & Decision Frameworks: Black-box reverse engineering research archive for the Jev decision model

  • jev-by-example - SDK & Decision Frameworks: Ten runnable Jev examples for agent decisions: memory conflicts, tool-result checks, recovery, context selection, and handoffs. JavaScript, zero dependencies.

  • jev-doom - SDK & Decision Frameworks: Watch Jev play Freedoom in a local dashboard. TypeSafe direct and Vercel AI Gateway, inspectable decisions, and bounded spending.

  • jev-go - SDK & Decision Frameworks: A small unofficial Go SDK for Jev calls and model listing.

  • jev-phone - SDK & Decision Frameworks: Drive a phone with a model that never writes a word. TypeSafe's Jev picks each action, phone-use runs it on iOS and Android.

  • jev-starter - SDK & Decision Frameworks: TypeScript patterns for thresholds, fallbacks, human review and evaluation on top of the TypeSafe SDK.

  • typesafe-ai-jev-example - SDK & Decision Frameworks: This repository provides six runnable Python demos and four notes covering TypeSafe Jev primitives and composition patterns, with offline mock mode and committed live samples from jev-1.13.0.

  • typesafe-go - SDK & Decision Frameworks: A TypeSafe System One client that uses only the Go standard library to send Jev questions and read structured answers.

  • typesafe-go - SDK & Decision Frameworks: An unofficial Go SDK with typed answers, retries and context cancellation.

  • typesafe-rs - SDK & Decision Frameworks: A community Rust Jev client with async requests, an optional blocking interface and local mock testing.

  • claude-jev-mod - SDK & Decision Frameworks: Typed decisions in Claude Code: adds $.jev over TypeSafe's Jev, through OpenRouter, Vercel AI Gateway, Cloudflare Workers AI, LiteLLM or the TypeSafe API.

  • jev_playground - SDK & Decision Frameworks: Jev answers typed questions about a state with calibrated numbers. This repo is where we find out which of those numbers deserve to drive code, and where a regex or a constant does the job better.

  • jev-go-sdk - SDK & Decision Frameworks: Dependency-free Go client for TypeSafe AI's System One API and the Jev model

  • jev-is-not-odd - SDK & Decision Frameworks: A probabilistic, AI-powered utility to determine if a number is not odd (or not even) using TypeSafe's Jev model and the Vercel AI SDK.

  • jev-lab - SDK & Decision Frameworks: This repository provides a single-page classifier demo that sends text with choice questions to Jev and displays the request JSON, probability distribution, confidence, latency, and token usage.

  • jev-lab - SDK & Decision Frameworks: Hands-on research lab for TypeSafe's Jev (System One model): reproducible benchmarks of Noul/Choice/Score primitives, confidence gating, fan-out latency, agent control — plus a living audit of the Jev ecosystem.

  • jev-labs - SDK & Decision Frameworks: Never confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong verdicts. Film, code, and every captured call.

  • jev-msw - SDK & Decision Frameworks: Mock Jev API decisions with MSW for deterministic tests without real API calls or credits.

  • jev-swap - SDK & Decision Frameworks: Find the LLM calls in your codebase that are really decisions, see what they'd save on TypeSafe Jev, and prove it on live traffic before you swap. TypeScript, JavaScript, Python.

  • jev-tab-order - SDK & Decision Frameworks: **Organize the entire window with a single Jev API request.** Grouping and ordering decisions are evaluated together, regardless of the number of tabs.

  • jev-triage - SDK & Decision Frameworks: Automated issue & PR triage for open-source maintainers, powered by Jev (TypeSafe AI).

  • jevcore - SDK & Decision Frameworks: A judgement primitive for TypeSafe **System One / Jev** — ask N things × K typed questions in bounded, cheap, fail-open requests — plus the **JEV harness**, the closed boundary in code that makes a Jev answer safe to consume. Standard library only.

  • jevinf - SDK & Decision Frameworks: An inference engine for decision models of the Jev kind: each candidate path runs as segmented forwards with prefix reuse, and the Jev wire contract is served on top. NanoJev is the backend wired up today.

  • jevish - SDK & Decision Frameworks: Every mode auto-curries when called with only the patterns:

  • jevpolicy - SDK & Decision Frameworks: JevPolicy is an open-source TypeScript decision runtime that turns probabilistic judgments from Jev, accessed through Vercel AI Gateway, into versioned, deterministic, replayable, observable application decisions.

  • Jevs-Garage - SDK & Decision Frameworks: System One turns unstructured or structured state into fast probabilistic judgments. Instead of asking for free-form prose, these demos ask `Choice`, `Score`, and `Noul` questions and receive typed values with uncertainty that application code can reason about.

  • typesafe-ai-ruby - SDK & Decision Frameworks: A stdlib-only Ruby client that sends Choice, Score, or Noul questions to TypeSafe System One.

  • typesafe-sdk-swift - SDK & Decision Frameworks: An experimental Swift SDK for TypeSafe using Swift Package Manager, Swift concurrency, and URLSession.

  • TypeSafeSDK - SDK & Decision Frameworks: An unofficial .NET client that POSTs state and typed questions to TypeSafe /v1/systemone. The parent repo also contains unrelated SDK dumps.

  • langchain - SDK Integrations: An optional Jev classifier integration for Python LangChain workflows.

  • pydantic-ai - SDK Integrations: An optional TypeSafe provider and Jev model integration for Pydantic AI.

  • ax - SDK Integrations: Ax provides a TypeSafe integration for boolean or finite-class signatures and native Jev answers.

  • ruby_llm-typesafe - SDK Integrations: A TypeSafe provider for RubyLLM 2 that exposes Jev’s three judgment types through structured output.

  • jev-resilience - SDK Integrations: A Spring WebFlux integration that detects error messages hidden in HTTP 200 response bodies.

  • laya-mlx - Local runtime: independent MLX port of the Laya checkpoints that runs typed decisions natively on Apple Silicon — 13.4 ms median end-to-end per short English decision, 7.4 ms with the multilingual checkpoint, and zero output tokens, with no PyTorch, Transformers runtime, or cloud API.

  • laya-Ascend - Local runtime: Ascend NPU fork of the Laya checkpoints that answers the same Choice, Score and Noul questions on Huawei 910B hardware — 37–47 ms median for a four-question request, 33.8x–70.9x faster than the same request on a single container CPU thread, with an output-equivalent SDPA decision head that avoids torch_npu's CPU fallback on aten::_transformer_encoder_layer_fwd.

  • jev-trust - Python ecosystem: trust middleware for the Jev API that logs every typed decision, measures calibration in your own domain from outcomes you record (accuracy, Brier, top-label ECE, C = 1 − ECE), annotates each answer with its measured effective confidence, fires overconfidence alerts, and signs the evidence (ed25519) for independent recomputation.

  • JarvisCore - Agent frameworks: Python multi-agent runtime that ships Jev natively from 1.12, where agents ask typed Choice, Score and Noul questions through a decision client separate from the text model, the Kernel picks a specialist subagent by Choice, and each retrieved RAG passage is withheld from the generating model when its prompt-injection Noul exceeds 0.70.

  • hunch (carldaws) - Ruby ecosystem: turns judgment calls into control flow — if Hunch.likely?("fraudulent", given: order) reads like plain Ruby but branches on a typed Jev answer, with pick for Choice, rate for Score, and graded predicates from possibly? to almost_certainly?; a TypeScript port offers the same interface.

  • Early experimentation using Jev to rethink harness UX - Harness integration: an agent platform wires Jev into its LLM harness as a callable tool for search, approvals and context, reporting 2,000 expense reports categorized in 20 seconds for five cents.

  • jev-mcp (burnigtm) agent: Multi - MCP ecosystem: server that puts Jev into the coding loop for Cursor, Codex, and any MCP client, with 20 test files behind it.

  • jev-architect - Design skill: finds, designs, and evaluates Jev decision loops, packaged as a skill with references on decision design and delivery.

  • Building a Harness with Jev - Framework guide: LangChain's walkthrough of wiring Jev into an agent harness as the decision layer, from a team that then published its own evaluation of Jev as a judge.

  • system-one-adapter-python - Python ecosystem: TypeSafe AI's official open-source drop-in adapter for running and benchmarking Jev System One decision evaluations across OpenAI- and Anthropic-compatible LLM APIs.

  • neurolink - Provider abstraction: the pipe layer of an AI nervous system — Juspay's TypeScript interface connecting provider neurons to an application, with decide as a first-class inference type alongside generate and stream.

  • jev-spring-boot-starter - Java ecosystem: Spring Boot 4 starter that puts Jev behind Spring MVC and RestClient.

  • mysql-ailike - Database filtering: MySQL plugin that filters rows by a natural-language predicate instead of a literal one, powered by Jev.

  • FastJev - Local runtime: self-hosted Python SDK and System One-compatible API for runtime-defined Choice, Boolean, and Score decisions on pinned open models across Torch, vLLM, MLX, llama.cpp, and WebGPU, with committed row-level benchmarks and checksums.

  • Search with Jev and Milvus - Search engineering: nine runnable notebooks combine Gemini embeddings and Milvus retrieval with Jev Noul and Choice judgments, while Python applies ranking, filtering, routing, and stopping policies to synthetic examples.

  • jeff (logan-markewich) - Self-hosted runtimes: self-hosted drop-in replacement for TypeSafe Jev powered by GliFormer, exposing native Choice, Score, and Noul decision endpoints without cloud API dependencies.

  • CloJev - Clojure ecosystem: unofficial portable Clojure SDK for System One, so Clojure applications can put typed questions to Jev without a Java interop layer.

  • spring-ai-typesafe type: library - Java ecosystem: Java SDK for TypeSafe AI's Jev API and Spring AI integration, providing typed decisions for evaluation as a judge, guardrails, and RAG post-processing.

  • JevFlow type: library - TypeScript ecosystem: composes Jev Noul, Score, and Choice decisions into deterministic threshold workflows that batch into a single systemOne call and return an ordered, explainable action set instead of side effects, with matched rules recording the actual value behind each action and a mock provider so policy tests run without an API key.

  • Jev AI Tools type: hosted - Developer education: hosts six bounded recipes plus a custom builder for AI SDK Choice, Score, and Boolean evaluations; recipes display answer probabilities separately from provider confidence and use deterministic local thresholds to pause uncertain routes for review, while every configuration exports as TypeScript.

  • jev-foundation-models - Apple platforms: a Swift 6 bridge that runs Jev decisions through Apple's Foundation Models on device, with a protocol-based model interface and six tests.

  • cu-Jev - Inference engine: a CUDA-native implementation of the Jev System One API that keeps decisions GPU-resident, shipping a Starfighter demo and a benchmark script.

  • jevcache - Cost control: memoizes Jev-class decisions so a repeated question is served from cache instead of a new call, keeping repeats deterministic and free.

  • Qwev type: self-hosted - Local inference: turns dense Qwen3 and Qwen3.5 checkpoints into a training-free Jev-style Noul, Choice, and Score service that shares one state prefill across isolated questions and, on its included 27-question Qwen3.5-9B/A100 fixture, reports 0.500 s versus 13.554 s for generated JSON.

  • jev-sdk-go type: library - Go ecosystem: dependency-free Go 1.24+ client for Jev Noul, Choice and Score questions that reads Choice and Score answers back as the caller's own types, rejecting any label or level the question never offered, with retries, OpenRouter support, and eleven examples tested against an in-process fake of the API.

  • jevcompat type: cli - Interoperability: a 48-requirement spec of the POST /v1/systemone wire contract, each rule citing TypeSafe's docs, OpenAPI file or SDKs, and a suite that checks any Jev-compatible server against it (Choice probabilities keyed by option and summing to 1, Score equal to Σ i·p, 2–255 options, error shapes, answers that stay put when question ids or order change), finding 2 of the 8 most-starred open ports conformant, with a reference mock that breaks each rule on purpose and a proxy that fixes what can be fixed.

  • typesafe-ai-php - PHP ecosystem: A modern TypeSafe AI client and SDK, with result classes and classic requests

  • Jev type: library - PowerShell ecosystem: PowerShell module for building Jev Noul, Choice, and Score questions and returning named answers as pipeline-friendly properties.

  • djev-run - Serving: deploys DiffusionGemma-Jev behind a TypeSafe-compatible API on a Cloud Run GPU with snake, dino and tetris demos wired to the decision endpoint.

  • jev-symfony-bundle type: library - PHP / Symfony: Symfony bundle providing typed Jev clients, validation constraints (#[JevNoul], #[JevChoice]), Workflow guards, and WebProfiler panels.

  • JevT++ type: library - C++ integration: independent C++20 library with compile-time enum schemas, typed Choice/Noul/Score results and abstention, local Laya inference through ONNX Runtime or ggml, and an opt-in TypeSafe System One HTTP backend tested with mocks and loopback HTTP rather than live-provider calls.

  • Sim - Agent frameworks: open-source collaborative workspace for building, deploying, and monitoring AI agents featuring native TypeSafe System One evaluation and decision provider integration.

  • RubyLLM type: library - Ruby ecosystem: official Ruby gem connecting TypeSafe judgment models to RubyLLM with a native System One protocol for typed questions, probabilistic answers, and error normalization.

  • jev-style agent: Multi type: self-hosted - Local runtime: pip install "jev-style[torch]" (or [mlx] on Apple silicon) serves an open 0.8B Qwen3.5 decision model behind a System One-compatible /v1/systemone API that answers Choice, Score, and Noul questions with calibrated probabilities in one pass over inputs up to 25,600 tokens (0.15–0.2 s per short request with MLX on an M1 Max, after the first call), and ships a Claude Code guard that turns four Noul checks and a risk Score into allow, ask, or deny through code-owned thresholds.

  • grev type: cli - Developer tooling: grep, sort, cut and uniq that match by meaning — grev 'is a vegan meal' menu.txt keeps the lines whose Jev Noul clears 0.5, sibling filters route by Choice and rank by Score, and the output is always your own input, never generated text.

  • decision-gate type: library - Cost and rate control: npm library that every Jev request in a loop goes through, which waits for room under 80% of the account's requests-per-minute and tokens-per-second limits, pauses every caller sharing the account, across processes, for the server's Retry-After delay when the service answers 429, refuses any request that would pass a per-key daily spend ceiling (USD 0.20 by default), and keeps an opt-in cache of answer probabilities under caller-chosen keys, so a repeated question over the same state is not paid for twice.

  • ollaya type: self-hosted - Local runtime: serves open decision models behind a wire-identical /v1/systemone endpoint, so an existing Jev client only has to point TYPESAFE_BASE_URL at the daemon, and reports its recommended model at 0.722 accuracy against Jev's 0.738 on typed decisions.

  • Tiltmeter type: proxy - Monitoring: drop-in /v1/systemone proxy and Pydantic AI client that records every Jev answer's probabilities and alerts, without labels, when jev-latest switches versions, a question's answers drift (chi-square-tested PSI), answers crowd a decision threshold, or estimated accuracy falls.

  • DecisionKit type: library - .NET ecosystem: provider-independent .NET decision engine whose domain package holds no Jev URL, header or DTO, mapping Choice, Score and Noul questions onto POST /v1/systemone from a separate provider package, with a runnable ASP.NET ticket-triage sample that picks the owning team and escalates to a human at a normalized Score of 0.8, and 1,096 tests across net8.0 and net10.0 that run with no HTTP.

  • intern-decision-mlx type: self-hosted - Local runtime: serves Shanghai AI Lab's open Intern-Decision-0.8B vision decision model on Apple Silicon behind a /v1/systemone-shaped endpoint that also takes screenshots, so an agent can threshold the calibrated Choice, Score and Noul answers and escalate the rest — 0.9 s per 1080p screenshot downscaled to 1024 px on an 8 GB M2 MacBook Air, and the same answers as the lab's PyTorch reference on 67 of 67 fields.

Game & Simulation

Source file: categories/game-simulation.md

  • typesafe-mario - Gaming: TypeSafe/Jev agent that plays Super Mario Bros. from structured emulator state, choosing each action from emulator-derived features.

  • jev-drone - Robotics simulation: camera-only autonomous drone in MuJoCo that puts a Jev judgment model in the control loop at 2.5 Hz.

  • tsai-sc - Gaming: drives original StarCraft shareware through keyboard and mouse with Jev action probabilities recorded per decision.

  • jev-plays-pokemon - Gaming: reads Pokémon Red game state as text, answers typed questions each turn, and lets deterministic code turn the answers into moves.

  • typesafe-jev-drone-demo - Simulation: Three.js drone simulator with a Python backend where Jev drives the navigation decisions.

  • typesafe-playground - Interactive playground: small Jev experiments that put the decision on screen, from routing a support message to steering a car in a 3D world.

  • PlayJev - Gaming: open 0.8B vision-language model that reads one 448 px game frame, returns a probability over the moves the game lists in a single forward pass with no generated text, and hands its low-confidence steps to a search program, across ten browser games.

  • jev-plays-pokemon-red - Gaming: Pokemon Red on PyBoy where deterministic code owns the route and arithmetic, Jev picks only at branches, and every battle turn's faint prediction is scored by Brier against RAM state.

  • jevpilot - High-Frequency / Games: A browser driving simulator where Jev chooses among locally generated paths and speeds.

  • jevk5 - High-Frequency / Games: An open model that answers the same typed questions Jev answers — yes/no, choice, score — in one forward pass with zero generated tokens (~13 ms on an H100), plus a head-to-head harness that puts Jev and the open model on the identical board and question so the two can be compared directly. It serves TypeSafe's `/v1/systemone` wire format, so a Jev client can point at it unchanged.

  • laya-vs-jev - High-Frequency / Games: Laya vs Jev: local MLX and hosted AI decisions playing T-Rex side by side, with live metrics and replay recording

  • jev-libero - High-Frequency / Games: Two LIBERO tasks, one control engine. Each demo loads its own JSON task definition. Videos follow simulation time, with decision and physics-preview waiting omitted.

  • RoboJEV - High-Frequency / Games: RoboJEV is a small, inspectable robotics laboratory. JEV receives **structured simulator state, not images**, selects an immediate intent, then selects X/Y/Z directions and a gripper command. A Cartesian controller executes the action using real MuJoCo contacts. Each task has independent physical success checks; model answers cannot declare success.

  • OneVOneJev - High-Frequency / Games: A browser-based 1v1 shooter where Jev reads structured match state and chooses movement, aim, and firing.

  • laya-vs-jev-arena - High-Frequency / Games: Laya (open source, local) vs TypeSafe Jev (API): two AI models race in Snake and fight in a Mortal-Kombat-style arena. Every move is a real model decision.

  • jev-doom-agent - High-Frequency / Games: A browser Doom experiment comparing Jev-controlled players from the same initial state.

  • typesafe-snake - High-Frequency / Games: Snake autoplayer driven by TypeSafe Jev, executing one discrete System One decision per tick with code-enforced legal moves.

  • live-jev - High-Frequency / Games: A browser-based top-down driving simulator using Jev for lane and speed choices, with an optional chat-model comparison.

  • jev-reflex-autonomy-lab - High-Frequency / Games: Multi-drone autonomy lab demonstrating TypeSafe Jev reflex decisions with optional System 2 strategy guidance.

  • jev-askable-arm - High-Frequency / Games: Uses Jev to chain predefined skills for English-language goals in a ManiSkill robot-arm simulation.

  • jevscape - High-Frequency / Games: A RuneBench extension using Jev and a bounded rs-sdk action catalog for RuneScape tasks.

  • jevtown - High-Frequency / Games: A check costs from half a cent (a text that dies in the first wave) to ten cents (one that reaches all 10,000), and takes from 3 seconds to a minute. The interface comes in Ukrainian and English, and so do the personas: a text is read by the crowd that speaks its language, 10,000 Ukrainians or 10,000 English speakers, so there is nothing to choose.

  • heist-one - High-Frequency / Games: Observable browser stealth game where Jev makes typed guard judgments while deterministic code owns the physics world.

  • jev_deep_rl - High-Frequency / Games: This project evaluates a fixed model. It records rewards and decisions without training or updating model weights. A seeded random policy provides a local baseline.

  • JevBird - High-Frequency / Games: A Python Flappy Bird game where code simulates candidate routes and Jev picks one.

  • doom-jev - High-Frequency / Games: A ViZDoom Agent that uses Jev to choose movement, targets and firing from structured game state.

  • jev-little-airways - High-Frequency / Games: An island-airport simulator using Jev for routes, yielding, emergency broadcasts and landing order.

  • soupbase - High-Frequency / Games: Soupbase is a bilingual Chinese-English Turtle Soup game where Jev judges player questions and reconstructions, and the app checks structured Choice results and confidence to decide clearance.

  • jev_vampire_survivors - High-Frequency / Games: TypeSafe's Jev model plays Vampire Survivors on Steam: BepInEx plugin + Python brain + live decision dashboard. Native Linux only.

  • jev-robotics-demo - High-Frequency / Games: A MuJoCo arm demo where local code proposes candidate moves and Jev chooses the target, grasp or release, and completion.

  • jev-arena-nanojev - High-Frequency / Games: Jev Arena is a fully local grid tactical game arena where NanoJev, rule agents and search algorithms make per-step move, attack, shoot, heal, dash and environment-interaction decisions across multiple levels with a Chinese Pygame interface.

  • jev-flappy-bird - High-Frequency / Games: A live demo of TypeSafe's Jev model playing Flappy Bird, one flap-or-wait decision at a time.

  • jev-gamepilot - High-Frequency / Games: **Jev-GamePilot** is a universal autonomous AI gaming agent powered by **Laya (local sub-30ms System One inference)** and **TypeSafe's Jev System One** (`Choice`, `Score`, `Noul`). It captures real-time gameplay at 60+ FPS, fuses instant local reflexes with high-level strategic reasoning, and executes physical hardware inputs across Windows PC games and connected Android phones.

  • jev-gpt - High-Frequency / Games: Cascaded Choice questions that make Jev pick the next word from a word tree instead of generating text.

  • jev_fsd - High-Frequency / Games: **An AI model drives a car through a real city, and you can watch every decision it makes.**

  • jev-market-reflex - High-Frequency / Games: Fast typed AI decisions on live crypto markets using TypeSafe AI Jev.

  • jev-play-ping-pong - High-Frequency / Games: Uses Jev to choose serve direction, return angle and pace in a browser table-tennis game.

  • jev-rl - High-Frequency / Games: JEV Reinforcement Learning: four classic games trained with JEV-powered rewards, reproducible experiments and checkpoint replays.

  • jevarena - High-Frequency / Games: Two Jev Agents play Snake in side-by-side browser panes with visible per-step choices.

  • snake-jev - High-Frequency / Games: Real-time Snake game driven by parallel Jev assessments, deciding optimal turns in a single API call per tick.

  • tsai-civ2 - High-Frequency / Games: An experimental harness where TypeSafe Jev plays classic Civilization II in a browser, computing live action probability distributions.

  • jev-broadcast-lab - High-Frequency / Games: A Jev experiment workbench centered on chess, with additional classification and matching exercises.

  • jev-chess - High-Frequency / Games: Chess moves, evaluations, persona opponents, and game classification using TypeSafe AI System One models. Resolves natural language move intents into legal moves, evaluates positional sharpness and king risk in parallel, and powers historical persona opponents (Tal, Capablanca, Petrosian).

  • jev-flappy-bird - High-Frequency / Games: Jev learns to play flappy-bird game with physics based context and without it

  • jevTrader - High-Frequency / Games: A High Frecuncy Trader made in Rust using Jev as a decision maker.

  • mk-jev-fly-brain - High-Frequency / Games: Compares a fly-connectome spiking simulation, Jev and rule policies in the mk.js fighting game.

  • can-jev-bayes - High-Frequency / Games: How well can Jev make sequential decisions under uncertainty, and how can Bayesian methods help it learn and act more effectively?

  • jev-claim-vs-measured - High-Frequency / Games: A post with ~400k views says TypeSafe's **Jev** is the fastest AI model ever built for trading, makes calibrated buy/sell decisions in under 100 ms, and shows how to build an HFT system on it. The article behind it contains no backtest, no P&L and no hit rate. So I ran the tests: on the raw tape at one decision per second, and at 15–60 minute horizons. It cost **$0.97** of API credit. Everything needed to check me is in this repository.

  • jev-clash-royale-test - High-Frequency / Games: A Clash Royale-style sandbox whose Jev bot decides play-or-hold, card, lane, and depth in one System One call.

  • jev-experiments - High-Frequency / Games: Uses Jev to play Chrome Dino and a local shooter arena while Python executes structured decisions.

  • jev-factorio-agent - High-Frequency / Games: Jev picks what, code owns how - a System One Factorio agent driven by TypeSafe's Jev on FLE

  • jev-practice-speed - High-Frequency / Games: A WebGL demo where you play the card game Speed against a CPU whose brain is TypeSafe AI's Jev. The whole point of the app is to measure and show Jev's decision speed and decision accuracy in real time.

  • jev-synthetic-survey - High-Frequency / Games: New to synthetic survey respondents? [Start here](#new-to-this-start-here). For the raw runs, the scored reports and the code, see [where to go](#where-to-go).

  • jev-table-tennis - High-Frequency / Games: Table tennis vs. TypeSafe's Jev (System One). Every paddle move on the right is a live model decision — no local prediction, just a lookup table and a servo.

  • typesafe-jev-decision-studio - High-Frequency / Games: Fast, calibrated System One decision platform powered by TypeSafe Jev via OpenRouter. Sub-second logprob scoring, transfer curves, zero hallucinations.

  • typesafe-jev-traffic-demo - High-Frequency / Games: This is a **simulation**. It is not connected to, and cannot control, any real traffic signal — Hong Kong's Transport Department publishes no write API for that, only a read-only feed of sensor data. Everything downstream of that feed (the phase timing, the amber/all-red clearance, the safety limits) runs entirely in this process's own memory.

  • [jev-torneo-animales](https:

Truncated — view the full README on GitHub.

awesome
awesome-ai
awesome-jev
awesome-list
awesome-lists
awesome-readme
awesome-resources
calibrated-probabilities
jev
jev-ai
jev-alternative
jev-api
jevbench
jev-model
laya
llm
llm-evaluation
llm-inference
llm-tools
vmodal

Significant stargazers

Leo Chiu

65 followers · starred Sep 2026