feder-cr/jev

jevos is an open-source alternative to Jev for yes/no decisions that runs on your laptop.

C++

1,126

28 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

The fastest model ever made ten times faster than JEV (r/SideProject)

Honestly got sick of shipping every little yes/no decision off to the cloud. Felt ridiculous. Full LLMs for that? Come on. So I threw together jevos. You dump some text and a yes/no question at it, it does one forward pass and spits back the probability of yes. No fancy text generation, just the…

10

Sep 30, 2026

README

jevos

Yes/no decisions on a laptop CPU in 25–110 ms. Send a text and a yes/no question, get back P(yes).

jevos playing a Chrome Dino-style game on the CPU, answering two yes/no questions per step (recording at 2× speed)

Benchmarks

Latency on a short request: jevos-v2 26 ms, Jev 344 ms, Laya 104 ms. On a long request: jevos-v2 112 ms, Jev 345 ms, Laya 449 ms. Accuracy on 2,000 yes/no questions from unseen policies: jevos-v2 0.810, Jev 0.927, Laya 0.489

jevos-v2 answers 80.3% of 999 hand-written yes/no questions correctly, against 75.8% for the first jevos, with the same size and speed.

Latency is the median of 10 requests through the HTTP API, after 3 warm-up requests, on an Intel Core Ultra 7 255H laptop with 16 threads, each request reading its text from scratch (--state-cache 0). By default the server keeps the texts it has read, so asking about the same text again takes 22 ms for the long request.

What each one does

jevos-v2JevLaya
Yes/no questions✓✓✓
Multiple choiceSoon✓✓
ScoresSoon✓✓
Runs onyour machinecloudyour machine
Costfreeper tokenfree
Context8,192 tokensnot stated512 tokens

Quickstart

From the release, download the archive for your system (jev-windows-x64.zip, jev-linux-x64.tar.gz or jev-macos-arm64.tar.gz) and the model, jevos-v2-openvino-int8.zip. Unpack the model into the jev folder, so that it sits in jev/model:

tar -xzf jev-linux-x64.tar.gz                 # Windows: unzip jev-windows-x64.zip
cd jev
unzip ../jevos-v2-openvino-int8.zip           # creates model/
./jev serve                                   # Windows: jev.exe serve
curl http://127.0.0.1:8017/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "jev-latest",
  "state": "I was charged twice for the same order.",
  "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}}}'
{
  "model": "jevos-v2",
  "answers": {"billing": {"type": "noul", "noul": 0.94}},
  "usage": {"input_tokens": 27, "output_tokens": 0}
}

jev runs on the CPU: jevos-v2 with 8-bit weights through OpenVINO, in one binary with no Python and no GPU. The same release has the model as GGUF files (jevos-v2-q4_k_m.gguf, jevos-v2-q8_0.gguf) for llama.cpp and the tools built on it.

API

The server speaks TypeSafe Jev's wire format, so code written for Jev's SDK works unchanged for yes/no questions.

POST /v1/systemone

FieldWhat it is
modeljev-latest (any jev-* name works) or the served model's name
statethe text to decide on: a string, or any JSON object or array
questionsone or more named questions, each {"type": "noul", "instructions": "…?"}

Every answer is noul, the probability that the answer is yes (0 to 1). Questions in the same request share the state, which is read once: the three questions below take about 66 ms together, against 49 ms for one of them alone, and 39 ms when the same state is asked about again.

{
  "model": "jev-latest",
  "state": {
    "item": "wireless mouse",
    "delivered": "5 days ago",
    "customer_message": "The box arrived empty. This is the second time!"
  },
  "questions": {
    "refund": {
      "type": "noul",
      "instructions": "Our policy refunds items reported missing within 30 days of delivery. Should this customer get a refund?"
    },
    "upset": {"type": "noul", "instructions": "Is the customer upset?"},
    "wrong_item": {"type": "noul", "instructions": "Does the customer say they received the wrong item?"}
  }
}
{
  "model": "jevos-v2",
  "answers": {
    "refund": {"type": "noul", "noul": 0.93},
    "upset": {"type": "noul", "noul": 0.83},
    "wrong_item": {"type": "noul", "noul": 0.04}
  },
  "usage": {"input_tokens": 95, "output_tokens": 0}
}
  • Put the rule in the question. If the decision depends on a policy, write it into instructions, as in refund above. Jev's optional criteria field is accepted but not read.
  • Yes/no only. choice and score questions are refused with a 422.
  • Timing. Every response carries a Server-Timing header: parsing, validation, tokenization, queue and inference times.

Other endpoints

EndpointReturns
GET /v1/modelsthe served model and its jev-latest alias
GET /health{"status": "ready", …} once the model is loaded

From Python

import requests

answer = requests.post("http://127.0.0.1:8017/v1/systemone", json={
    "model": "jev-latest",
    "state": "I was charged twice for the same order.",
    "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}},
}).json()

if answer["answers"]["billing"]["noul"] > 0.5:
    print("send to billing")

From the command line

jev decide answers one request file without starting a server. The file holds the same body as POST /v1/systemone, and the answer comes back in the same shape:

./jev decide request.json

--output answer.json writes the answer to a new file instead of printing it; an existing file is never overwritten. A request the server would refuse prints the same error body on stderr, with exit status 1.

Server options

OptionDefault
--model-dirmodel beside the binarythe model folder
--threadsall logical CPUsCPU threads; fewer if other heavy apps are running
--host, --port127.0.0.1, 8017where the server listens
--state-cache, --state-cache-tokens16, 8,192texts kept for later requests, how many and how many tokens in all; 0 turns it off
--batch-tokens384small requests arriving together are read in one model call while their tokens fit
--dynamic-quantization128activations in INT8 groups of this many values; 0 keeps them f32, slower

With JEV_API_KEY set, every call but /health needs Authorization: Bearer <key>. The server needs about 1 GB of memory once the model is loaded, up to 1.4 GB with its cache of recent texts full. /health reports the SHA-256 of each model file and a fingerprint of them all (model_files, fingerprint), so a logged decision can be tied to the exact model that made it.

Build from source

python -m pip install -r requirements.txt     # OpenVINO's SDK, CMake, Ninja, the tests' packages
python scripts/build.py                       # dist/jev; on Windows, from a Visual Studio developer prompt
python tests/check.py                         # with the model in dist/jev/model

A C++17 compiler is the only other requirement; CMake fetches llama.cpp (the tokenizer) itself. export/export_openvino.py is how the model folder was made from the trained weights.

Guides

The wiki has the long version: zero-shot text classification with yes/no questions, LLM as a judge on a CPU, policy decisions, and what we measured about the model, including why a small LLM says yes when the answer is no.

Credits

Built together with Loris Salsi (@LosaLosSantos).

binary-classification
classification
cpu-inference
decision-model
edge-ai
fastapi
gguf
jev
llama-cpp
llm
llm-inference
local-llm
minicpm
nlp
offline
on-device-ai
python
self-hosted
text-classification
yes-no

feder-cr/jev

jevos is an open-source alternative to Jev for yes/no decisions that runs on your laptop.

C++

1,126

28 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

The fastest model ever made ten times faster than JEV (r/SideProject)

Honestly got sick of shipping every little yes/no decision off to the cloud. Felt ridiculous. Full LLMs for that? Come on. So I threw together jevos. You dump some text and a yes/no question at it, it does one forward pass and spits back the probability of yes. No fancy text generation, just the…

10

Sep 30, 2026

README

jevos

Yes/no decisions on a laptop CPU in 25–110 ms. Send a text and a yes/no question, get back P(yes).

jevos playing a Chrome Dino-style game on the CPU, answering two yes/no questions per step (recording at 2× speed)

Benchmarks

Latency on a short request: jevos-v2 26 ms, Jev 344 ms, Laya 104 ms. On a long request: jevos-v2 112 ms, Jev 345 ms, Laya 449 ms. Accuracy on 2,000 yes/no questions from unseen policies: jevos-v2 0.810, Jev 0.927, Laya 0.489

jevos-v2 answers 80.3% of 999 hand-written yes/no questions correctly, against 75.8% for the first jevos, with the same size and speed.

Latency is the median of 10 requests through the HTTP API, after 3 warm-up requests, on an Intel Core Ultra 7 255H laptop with 16 threads, each request reading its text from scratch (--state-cache 0). By default the server keeps the texts it has read, so asking about the same text again takes 22 ms for the long request.

What each one does

jevos-v2JevLaya
Yes/no questions✓✓✓
Multiple choiceSoon✓✓
ScoresSoon✓✓
Runs onyour machinecloudyour machine
Costfreeper tokenfree
Context8,192 tokensnot stated512 tokens

Quickstart

From the release, download the archive for your system (jev-windows-x64.zip, jev-linux-x64.tar.gz or jev-macos-arm64.tar.gz) and the model, jevos-v2-openvino-int8.zip. Unpack the model into the jev folder, so that it sits in jev/model:

tar -xzf jev-linux-x64.tar.gz                 # Windows: unzip jev-windows-x64.zip
cd jev
unzip ../jevos-v2-openvino-int8.zip           # creates model/
./jev serve                                   # Windows: jev.exe serve
curl http://127.0.0.1:8017/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "jev-latest",
  "state": "I was charged twice for the same order.",
  "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}}}'
{
  "model": "jevos-v2",
  "answers": {"billing": {"type": "noul", "noul": 0.94}},
  "usage": {"input_tokens": 27, "output_tokens": 0}
}

jev runs on the CPU: jevos-v2 with 8-bit weights through OpenVINO, in one binary with no Python and no GPU. The same release has the model as GGUF files (jevos-v2-q4_k_m.gguf, jevos-v2-q8_0.gguf) for llama.cpp and the tools built on it.

API

The server speaks TypeSafe Jev's wire format, so code written for Jev's SDK works unchanged for yes/no questions.

POST /v1/systemone

FieldWhat it is
modeljev-latest (any jev-* name works) or the served model's name
statethe text to decide on: a string, or any JSON object or array
questionsone or more named questions, each {"type": "noul", "instructions": "…?"}

Every answer is noul, the probability that the answer is yes (0 to 1). Questions in the same request share the state, which is read once: the three questions below take about 66 ms together, against 49 ms for one of them alone, and 39 ms when the same state is asked about again.

{
  "model": "jev-latest",
  "state": {
    "item": "wireless mouse",
    "delivered": "5 days ago",
    "customer_message": "The box arrived empty. This is the second time!"
  },
  "questions": {
    "refund": {
      "type": "noul",
      "instructions": "Our policy refunds items reported missing within 30 days of delivery. Should this customer get a refund?"
    },
    "upset": {"type": "noul", "instructions": "Is the customer upset?"},
    "wrong_item": {"type": "noul", "instructions": "Does the customer say they received the wrong item?"}
  }
}
{
  "model": "jevos-v2",
  "answers": {
    "refund": {"type": "noul", "noul": 0.93},
    "upset": {"type": "noul", "noul": 0.83},
    "wrong_item": {"type": "noul", "noul": 0.04}
  },
  "usage": {"input_tokens": 95, "output_tokens": 0}
}
  • Put the rule in the question. If the decision depends on a policy, write it into instructions, as in refund above. Jev's optional criteria field is accepted but not read.
  • Yes/no only. choice and score questions are refused with a 422.
  • Timing. Every response carries a Server-Timing header: parsing, validation, tokenization, queue and inference times.

Other endpoints

EndpointReturns
GET /v1/modelsthe served model and its jev-latest alias
GET /health{"status": "ready", …} once the model is loaded

From Python

import requests

answer = requests.post("http://127.0.0.1:8017/v1/systemone", json={
    "model": "jev-latest",
    "state": "I was charged twice for the same order.",
    "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}},
}).json()

if answer["answers"]["billing"]["noul"] > 0.5:
    print("send to billing")

From the command line

jev decide answers one request file without starting a server. The file holds the same body as POST /v1/systemone, and the answer comes back in the same shape:

./jev decide request.json

--output answer.json writes the answer to a new file instead of printing it; an existing file is never overwritten. A request the server would refuse prints the same error body on stderr, with exit status 1.

Server options

OptionDefault
--model-dirmodel beside the binarythe model folder
--threadsall logical CPUsCPU threads; fewer if other heavy apps are running
--host, --port127.0.0.1, 8017where the server listens
--state-cache, --state-cache-tokens16, 8,192texts kept for later requests, how many and how many tokens in all; 0 turns it off
--batch-tokens384small requests arriving together are read in one model call while their tokens fit
--dynamic-quantization128activations in INT8 groups of this many values; 0 keeps them f32, slower

With JEV_API_KEY set, every call but /health needs Authorization: Bearer <key>. The server needs about 1 GB of memory once the model is loaded, up to 1.4 GB with its cache of recent texts full. /health reports the SHA-256 of each model file and a fingerprint of them all (model_files, fingerprint), so a logged decision can be tied to the exact model that made it.

Build from source

python -m pip install -r requirements.txt     # OpenVINO's SDK, CMake, Ninja, the tests' packages
python scripts/build.py                       # dist/jev; on Windows, from a Visual Studio developer prompt
python tests/check.py                         # with the model in dist/jev/model

A C++17 compiler is the only other requirement; CMake fetches llama.cpp (the tokenizer) itself. export/export_openvino.py is how the model folder was made from the trained weights.

Guides

The wiki has the long version: zero-shot text classification with yes/no questions, LLM as a judge on a CPU, policy decisions, and what we measured about the model, including why a small LLM says yes when the answer is no.

Credits

Built together with Loris Salsi (@LosaLosSantos).

binary-classification
classification
cpu-inference
decision-model
edge-ai
fastapi
gguf
jev
llama-cpp
llm
llm-inference
local-llm
minicpm
nlp
offline
on-device-ai
python
self-hosted
text-classification
yes-no

Languages

C++

54.8%

Python

43.7%

CMake

1.5%