ghraibeh/jev-saver

Local semantic memory in front of TypeSafe Jev: answers repeated decisions on CPU in ~30 ms, asks Jev only when unsure, and remembers the answer.

Python

0

9 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a local kNN cache in front of Jev. Numbers, caveats and two negative results inside (author here) (r/LocalLLaMA)

Disclosure first: I'm the author (Mahmoud, ghraibeh on GitHub). I'm not affiliated with TypeSafe AI. I know this sub is tired of Jev hype, so I'll lead with the limits. What it is: semantic caching plus nearest-neighbour voting. It isn't understanding and it isn't new. It's an MIT Python library…

0

Oct 4, 2026

README

jev-saver

Live demo: https://g-connect.space/jev-saver/ · Docs: https://g-connect.space/jev-saver/docs.html

jev-saver demo: a message answered locally, with the decision route on the right

A small local memory layer in front of TypeSafe Jev for repeated Choice decisions.

  • If a new input looks like inputs it has already seen, and those inputs agree, it answers locally. No Jev call is made.
  • Otherwise it asks Jev, returns Jev's answer, and stores it, so the same or a very similar input is answered locally next time.

It runs on CPU and keeps a small, saveable memory file.

What this is, honestly: semantic caching plus nearest-neighbour voting. It does not understand anything new. It only re-uses answers for inputs that look like inputs it has already seen. A rephrasing with different words often still goes to Jev. If your traffic rarely repeats, it will save little.

Screenshots

Answered locally. The message looked like ones already in memory, so Jev was not called (35 ms).

Answered locally

Asked Jev, then remembered. The memory was unsure (closest similarity 0.83), so it asked Jev (306 ms) and stored the answer. A repeat is answered locally.

Asked Jev

Bring your own key. The key is kept in the browser only, and the server does not store it.

Settings

Architecture
Architecture
Mobile
Mobile

Benchmarks page

Benchmarks

Why

Jev is already cheap (about $0.042 per 1M input tokens). For most people, money is not the main reason to use this. The main reasons are:

  • Latency: in our runs, a local answer took about 15–60 ms on CPU, compared with roughly 250–400 ms for a Jev round trip.
  • Rate limits: fewer requests count toward Jev's per-minute limits.
  • Offline / privacy: repeated inputs never leave the machine.

Install

pip install git+https://github.com/ghraibeh/jev-saver
export TYPESAFE_API_KEY=...

Use

from jev_saver import JevSaver

saver = JevSaver(
    labels={"billing": "Payment issues", "technical": "Bugs or errors", "account": "Login or profile"},
    instructions="Which team should handle this ticket?",
    model="jev-1.13.0",          # pin the version so stored answers stay consistent
)

d = saver.decide("I was charged twice this month")
print(d.source, d.choice, d.confidence)   # "jev" the first time, "local" for repeats

saver.warm_start(texts, labels)  # optional: pre-fill from labelled data you already have
saver.save("memory.npz")         # reuse later: saver.load("memory.npz")

Pre-fill memory (warm start)

If you already have labelled examples, load them before going live. No Jev calls are made. In the benchmark below, pre-filling raised calls saved from 53% to 87%.

import csv
rows = [r for r in csv.DictReader(open("tickets.csv")) if r["label"] in saver.labels]   # columns: text, label
saver.warm_start([r["text"] for r in rows], [r["label"] for r in rows])
saver.save("memory.npz")                       # once
saver = JevSaver(labels, instructions).load("memory.npz")   # in production

Sources you can use: tickets your team already tagged, logged Jev answers (keep only high-confidence ones), or a public dataset. You can also run decide() over a sample of real messages once with min_jev_confidence=0.9. Every label must exactly match a key in labels. Wrong labels get pre-filled too, so use only labels you trust.

Main knobs:

  • min_similarity: how close a stored input must be before the saver answers locally.
  • min_margin: how clearly the local vote must win.
  • exact_similarity: above this, the closest stored input is treated as the same input.
  • min_jev_confidence: do not store Jev answers that are below this confidence.

You can also pass a per-call key with decide(text, api_key=...), for example a key supplied by your end user. It is not stored.

Each Decision includes timings (in ms) for embed, search, jev and learn, plus learned.

Benchmark (offline, reproducible)

Run python benchmarks/banking77.py. It uses BANKING77 (77 intents) and needs no key.

The teacher in this benchmark is the dataset's gold label, standing in for Jev. That makes it a best case: real Jev is sometimes wrong, and the saver would store those mistakes. A live run against Jev has not been done yet.

Setup: 3,080 test messages in random order. The saver answers locally when it is sure; otherwise it asks the teacher and learns the answer. Default settings.

setupteacher calls savedlocal answers correctlocal mistakes
warm start (memory pre-filled from 10,003 training messages)87%97.6%65
cold start (empty memory, learns only from the stream)53%97.4%42

Cold start, by quarter of the stream: 31% → 54% → 65% → 64% answered locally.

Threshold sweep

Run it with python benchmarks/sweep.py.

startmin_similaritymin_marginsavedlocal correctmistakes
warm0.850.392%96.9%87
warm0.900.3 (default)87%97.6%65
warm0.940.664%98.9%22
cold0.850.368%95.3%99
cold0.900.3 (default)53%97.4%42
cold0.940.630%98.1%17

Raising either threshold means fewer local answers and fewer mistakes. Lowering them means more saving and more confident mistakes.

Limits we measured

  • No generalisation. In a separate test, the teacher labelled 548 new messages. Accuracy on other unseen messages stayed at 93.7%, the same as before. The saver remembers; it does not learn the concept.
  • Wrong but sure. About 2–3% of local answers are confidently wrong, and those are never sent to Jev.
  • English embedder. The default bge-small-en-v1.5 is weak on other languages. Swap the embedder= for a multilingual model if you need one.
  • Shifting answers. Stored answers are tied to the Jev version and to your labels. Clear the memory when either changes.
  • Jevstiller: the closest project. It trains a small head on frozen bge-small embeddings from Jev's recorded answers, and answers locally under a measured disagreement bound. jev-saver is simpler: there is no training step, just a nearest-neighbour vote with thresholds. It also has no statistical guarantee.

  • stuntd: a local proxy that learns from Jev and falls back to it.

  • decision-gate: handles Jev rate limits, with an optional answer cache.

  • GPTCache: semantic cache for LLMs.

  • jevcache: exact-match cache for Jev decisions.

  • FrugalGPT, arXiv:2305.05176: cascades that defer to a larger model only when needed.

License

MIT

ghraibeh/jev-saver

Local semantic memory in front of TypeSafe Jev: answers repeated decisions on CPU in ~30 ms, asks Jev only when unsure, and remembers the answer.

Python

0

9 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a local kNN cache in front of Jev. Numbers, caveats and two negative results inside (author here) (r/LocalLLaMA)

Disclosure first: I'm the author (Mahmoud, ghraibeh on GitHub). I'm not affiliated with TypeSafe AI. I know this sub is tired of Jev hype, so I'll lead with the limits. What it is: semantic caching plus nearest-neighbour voting. It isn't understanding and it isn't new. It's an MIT Python library…

0

Oct 4, 2026

README

jev-saver

Live demo: https://g-connect.space/jev-saver/ · Docs: https://g-connect.space/jev-saver/docs.html

jev-saver demo: a message answered locally, with the decision route on the right

A small local memory layer in front of TypeSafe Jev for repeated Choice decisions.

  • If a new input looks like inputs it has already seen, and those inputs agree, it answers locally. No Jev call is made.
  • Otherwise it asks Jev, returns Jev's answer, and stores it, so the same or a very similar input is answered locally next time.

It runs on CPU and keeps a small, saveable memory file.

What this is, honestly: semantic caching plus nearest-neighbour voting. It does not understand anything new. It only re-uses answers for inputs that look like inputs it has already seen. A rephrasing with different words often still goes to Jev. If your traffic rarely repeats, it will save little.

Screenshots

Answered locally. The message looked like ones already in memory, so Jev was not called (35 ms).

Answered locally

Asked Jev, then remembered. The memory was unsure (closest similarity 0.83), so it asked Jev (306 ms) and stored the answer. A repeat is answered locally.

Asked Jev

Bring your own key. The key is kept in the browser only, and the server does not store it.

Settings

Architecture
Architecture
Mobile
Mobile

Benchmarks page

Benchmarks

Why

Jev is already cheap (about $0.042 per 1M input tokens). For most people, money is not the main reason to use this. The main reasons are:

  • Latency: in our runs, a local answer took about 15–60 ms on CPU, compared with roughly 250–400 ms for a Jev round trip.
  • Rate limits: fewer requests count toward Jev's per-minute limits.
  • Offline / privacy: repeated inputs never leave the machine.

Install

pip install git+https://github.com/ghraibeh/jev-saver
export TYPESAFE_API_KEY=...

Use

from jev_saver import JevSaver

saver = JevSaver(
    labels={"billing": "Payment issues", "technical": "Bugs or errors", "account": "Login or profile"},
    instructions="Which team should handle this ticket?",
    model="jev-1.13.0",          # pin the version so stored answers stay consistent
)

d = saver.decide("I was charged twice this month")
print(d.source, d.choice, d.confidence)   # "jev" the first time, "local" for repeats

saver.warm_start(texts, labels)  # optional: pre-fill from labelled data you already have
saver.save("memory.npz")         # reuse later: saver.load("memory.npz")

Pre-fill memory (warm start)

If you already have labelled examples, load them before going live. No Jev calls are made. In the benchmark below, pre-filling raised calls saved from 53% to 87%.

import csv
rows = [r for r in csv.DictReader(open("tickets.csv")) if r["label"] in saver.labels]   # columns: text, label
saver.warm_start([r["text"] for r in rows], [r["label"] for r in rows])
saver.save("memory.npz")                       # once
saver = JevSaver(labels, instructions).load("memory.npz")   # in production

Sources you can use: tickets your team already tagged, logged Jev answers (keep only high-confidence ones), or a public dataset. You can also run decide() over a sample of real messages once with min_jev_confidence=0.9. Every label must exactly match a key in labels. Wrong labels get pre-filled too, so use only labels you trust.

Main knobs:

  • min_similarity: how close a stored input must be before the saver answers locally.
  • min_margin: how clearly the local vote must win.
  • exact_similarity: above this, the closest stored input is treated as the same input.
  • min_jev_confidence: do not store Jev answers that are below this confidence.

You can also pass a per-call key with decide(text, api_key=...), for example a key supplied by your end user. It is not stored.

Each Decision includes timings (in ms) for embed, search, jev and learn, plus learned.

Benchmark (offline, reproducible)

Run python benchmarks/banking77.py. It uses BANKING77 (77 intents) and needs no key.

The teacher in this benchmark is the dataset's gold label, standing in for Jev. That makes it a best case: real Jev is sometimes wrong, and the saver would store those mistakes. A live run against Jev has not been done yet.

Setup: 3,080 test messages in random order. The saver answers locally when it is sure; otherwise it asks the teacher and learns the answer. Default settings.

setupteacher calls savedlocal answers correctlocal mistakes
warm start (memory pre-filled from 10,003 training messages)87%97.6%65
cold start (empty memory, learns only from the stream)53%97.4%42

Cold start, by quarter of the stream: 31% → 54% → 65% → 64% answered locally.

Threshold sweep

Run it with python benchmarks/sweep.py.

startmin_similaritymin_marginsavedlocal correctmistakes
warm0.850.392%96.9%87
warm0.900.3 (default)87%97.6%65
warm0.940.664%98.9%22
cold0.850.368%95.3%99
cold0.900.3 (default)53%97.4%42
cold0.940.630%98.1%17

Raising either threshold means fewer local answers and fewer mistakes. Lowering them means more saving and more confident mistakes.

Limits we measured

  • No generalisation. In a separate test, the teacher labelled 548 new messages. Accuracy on other unseen messages stayed at 93.7%, the same as before. The saver remembers; it does not learn the concept.
  • Wrong but sure. About 2–3% of local answers are confidently wrong, and those are never sent to Jev.
  • English embedder. The default bge-small-en-v1.5 is weak on other languages. Swap the embedder= for a multilingual model if you need one.
  • Shifting answers. Stored answers are tied to the Jev version and to your labels. Clear the memory when either changes.
  • Jevstiller: the closest project. It trains a small head on frozen bge-small embeddings from Jev's recorded answers, and answers locally under a measured disagreement bound. jev-saver is simpler: there is no training step, just a nearest-neighbour vote with thresholds. It also has no statistical guarantee.

  • stuntd: a local proxy that learns from Jev and falls back to it.

  • decision-gate: handles Jev rate limits, with an optional answer cache.

  • GPTCache: semantic cache for LLMs.

  • jevcache: exact-match cache for Jev decisions.

  • FrugalGPT, arXiv:2305.05176: cascades that defer to a larger model only when needed.

License

MIT