Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
See the codeJev, distilled on the fly. Same call. Same answers. Your hardware.
Put Jevstiller in front of a repeated Jev classification call. At first every request still goes to Jev. From Jev's own answers — with their full probability distributions — it trains a small local model on your traffic, checks that the model agrees with Jev within a budget you set, and then answers most requests itself. Uncertain or novel input, and a permanent random audit slice, keep going to Jev.

You already call Jev. Run Jevstiller next to your service and point the SDK at it; nothing else changes:
docker run -d -p 8080:8080 -v jevstiller-data:/data ghcr.io/tomerglick57/jevstiller
export TYPESAFE_BASE_URL=http://localhost:8080 # your service keeps its own TYPESAFE_API_KEY
The recording above is examples/proxy_demo.py: 5,000 real customer messages through the unmodified SDK, 8 threads, against live Jev. The local model took over after about 4,000 requests. docker exec jevstiller jevstiller admin status <task> shows the audit agreement behind it (with -e JEVSTILLER_ADMIN_TOKEN=... on the container).
Jev is fast, cheap and typed. It is also ~300 ms away, per answer, at every load we tried (16 to 64 concurrent callers, up to 190 requests/s, p50 300 ms), and it is hosted: every classification is a network call to one vendor, under a published limit of 1,200 requests per minute that TypeSafe enforces at its own discretion. Jevstiller is for the workload where that hurts: decisions made one after another (an agent loop, a game tick, classify-then-act pipelines), latency budgets in milliseconds, boxes with no egress, or simply not wanting every classification to depend on one external API.
On Banking77 (77 customer-support intents), replayed against Jev's recorded answers, the local model answered 71.9% of held-out messages at 99.50% agreement with Jev (target 98%), taking over most traffic from about 5,000 messages on, and kept Jev's accuracy (78.5% vs Jev's 78.5% on the dataset's labels). The live run of 2026-09-25 gave 70.7% at 99.45%. A local answer takes ~15 ms on CPU (p50), about 20× faster than Jev; one CPU process answers ~130 messages/s, a GPU ~2,000/s.¹
Five public tasks, replayed through the loop with Jev's recorded answers as the teacher (bge-small on CPU, target agreement 98%). Answered locally and agreement are measured on 2,000 held-out rows against Jev; accuracy is against the dataset's own labels, for Jev alone and for the system (student where it answers, Jev elsewhere).
| Task | Answered locally | Agreement with Jev | Accuracy, Jev / system |
|---|---|---|---|
| Banking77, 77 intents | 71.9% | 99.50% | 78.5% / 78.5% |
| CLINC150, 150 intents † | 69.0% | 99.65% | 90.1% / 90.1% |
| AG News, 4 sections | 80.2% | 99.40% | 88.7% / 88.7% |
| TweetEval sentiment, 3 classes | 22.2% | 98.70% | 64.2% / 64.5% |
| TweetEval offensive, 2 classes | 24.1% | 98.80% | 73.8% / 74.2% |
† with rare_classes = "defer": Jev never used one of the task's labels.
Over 100 random splits across the 5 tasks, the calibrated threshold exceeded the 2% budget once; the usual point-estimate rule exceeded it on 6–12 of 20 splits per task.
Coverage tracks how consistent Jev itself is on a task, not how hard the task is: on the tweets Jev agrees with the dataset's labels only 64% and 74% of the time, so the student answers the confident quarter and forwards the rest, and the system's accuracy still matches Jev's. The threshold-rule comparison, what other targets buy, and the one-command reproduction: docs/benchmarks.md.
You set one number. Jevstiller returns the label Jev would have returned on at least that share of requests:
target_agreement = 0.98 → disagreement budget = 2% of all requests
The routing threshold is chosen on a held-out, IID calibration set so that, with 95% confidence, the share of requests the local model answers and gets different from Jev stays within the budget. It uses an exact finite-sample bound (Clopper–Pearson), testing candidate thresholds strictest-first. It is not tuned by eye, and it is re-verified forever on the audit channel. If the audit shows the contract is broken, everything falls back to Jev automatically.
If your code also acts on Jev's confidence, for example sending answers below 0.6 to review, set that number as the task's confidence_floor. The bound then also counts the requests Jev would have been less sure about than that, and local answers report at least that confidence. See configuration.
Agreement with Jev is not accuracy. If Jev is wrong, the student is wrong the same way. The status report says so next to every number. See DESIGN.md §2.
pip install jevstiller # everything: the proxy (`jevstiller serve`), the encoder, the Jev adapter
pip install "jevstiller[gpu]" # on a GPU machine: adds PyTorch, used automatically when CUDA is present
Python 3.10+. CPU works out of the box; a GPU only speeds up the encoder.
Run Jevstiller next to your services and point the Jev SDK at it:
docker run -d -p 8080:8080 -v jevstiller-data:/data ghcr.io/tomerglick57/jevstiller:0.4.0
# or: pip install jevstiller && jevstiller serve --config deploy/jevstiller.toml
export TYPESAFE_BASE_URL=http://jevstiller:8080 # in each calling service; nothing else changes
Services keep their own TYPESAFE_API_KEY; Jevstiller forwards it and never stores it.
Choice question becomes a task, keyed by its exact instructions, criteria and model, so services asking the same question share one local model.x-jevstiller-source: local tells you which), but only for keys Jev has accepted.Choice questions, other endpoints, and anything it does not understand.Performance of one process (16 vCPU, bge-small on CPU, docs/benchmarks.md):
Operations: a TOML config, an admin API and CLI (jevstiller admin ...), Prometheus /metrics, /readyz, JSON logs, jevstiller backup, and a Docker image and Kubernetes manifest. Docs: deploy, operations, the proxy, configuration, security, compatibility. Website: jevstiller.pages.dev.
As of September 2026:
stuntd train / stuntd enable.Without a key, python examples/quickstart_synthetic.py runs the whole loop on CPU in half a minute with a synthetic teacher standing in for Jev. With Jev (TYPESAFE_API_KEY in your environment):
from jevstiller import Task, Jevstiller, load_encoder
from jevstiller.teachers.jev import JevTeacher
task = Task(
name="support_router",
instructions="Which team should handle this customer message?",
classes={ # descriptions are sent to Jev verbatim — they are the spec
"billing": "Charges, invoices, refunds, payment methods",
"technical": "Bugs, errors, integrations, things not working",
"cancellation": "Wants to cancel, downgrade, or close the account",
"sales": "Pre-sales questions, plan comparison, quotes",
"other": "Anything that does not fit the categories above",
},
target_agreement=0.98,
)
js = Jevstiller(task, teacher=JevTeacher(), data_dir="./jevstiller-data", encoder=load_encoder("base"))
r = js.classify("Please cancel my subscription")
r.label # "cancellation"
r.confidence # 0.97
r.source # "teacher" at first — later "student:v9"
print(js.status().report())
Task: support_router version 89a7438c2f7d mode: cascade audit rate 2%
Production: student:v9 Shadow: -
Requests: 11,083 student 69.8% teacher 30.2%
Agreement with teacher (audit, n=415): 99.40% [98.46%, 99.84%] target 98% OK
note: agreement with the teacher is not accuracy.
bge-small/base/large), ONNX or PyTorch. Embeddings are stored, so retraining never re-encodes.student:vN directories; promote, roll back, or export() a standalone bundle.The full reasoning — including the five loop bugs the first real replay found and how they were fixed — is in DESIGN.md.
Alpha. Validated against live Jev (above).
Proxy: the drop-in proxy works end to end with the unmodified TypeSafe SDK.
Security: three security audits (12, 15 and 8 findings, all fixed: docs/security.md).
Deployment: it runs as a hardened container.
Tested under failure and load:
kill -9 during training.Results: docs/benchmarks.md.
Benchmarked on five public tasks against Jev's recorded answers (above); the full tables and the one-command reproduction are in docs/benchmarks.md.
See DEPLOYMENT_PLAN.md for what is done and what is next.
Documentation:
Not for: tasks with changing class lists (retrain from scratch), non-text input, or volumes too low to ever collect a few thousand examples.
Jev's answers for every Banking77 message ship with the repository, so the headline result replays without an API key:
git clone https://github.com/tomerglick57/Jevstiller && cd Jevstiller && pip install jevstiller
bash experiments/reproduce.sh # the Banking77 result above, on CPU, 10-15 minutes: 71.9% coverage at 99.50% agreement, deterministic
python experiments/run.py --dataset banking77 --teacher jev --encoder small # against live Jev (TYPESAFE_API_KEY)
python experiments/run.py --dataset banking77 --teacher oracle --encoder base # the dataset's labels as a perfect teacher
Apache 2.0 (see LICENSE and NOTICE). Jevstiller is an independent project and is not affiliated with, endorsed by, or supported by TypeSafe. "Jev" is their model; this tool only talks to its public API.
¹ Live run 2026-09-24: 11,083 replayed messages with 2,000 held out, bge-small on CPU; the GPU figure is from an oracle-teacher run with bge-base on an RTX 3090. Every number, with its command: docs/benchmarks.md.
Python
89.8%
Astro
3.5%
JavaScript
2.6%
CSS
2.0%
MDX
1.1%
Distill a repeated Jev classification task into a local model, on the fly — same answers, your hardware.
See the codeJev, distilled on the fly. Same call. Same answers. Your hardware.
Put Jevstiller in front of a repeated Jev classification call. At first every request still goes to Jev. From Jev's own answers — with their full probability distributions — it trains a small local model on your traffic, checks that the model agrees with Jev within a budget you set, and then answers most requests itself. Uncertain or novel input, and a permanent random audit slice, keep going to Jev.

You already call Jev. Run Jevstiller next to your service and point the SDK at it; nothing else changes:
docker run -d -p 8080:8080 -v jevstiller-data:/data ghcr.io/tomerglick57/jevstiller
export TYPESAFE_BASE_URL=http://localhost:8080 # your service keeps its own TYPESAFE_API_KEY
The recording above is examples/proxy_demo.py: 5,000 real customer messages through the unmodified SDK, 8 threads, against live Jev. The local model took over after about 4,000 requests. docker exec jevstiller jevstiller admin status <task> shows the audit agreement behind it (with -e JEVSTILLER_ADMIN_TOKEN=... on the container).
Jev is fast, cheap and typed. It is also ~300 ms away, per answer, at every load we tried (16 to 64 concurrent callers, up to 190 requests/s, p50 300 ms), and it is hosted: every classification is a network call to one vendor, under a published limit of 1,200 requests per minute that TypeSafe enforces at its own discretion. Jevstiller is for the workload where that hurts: decisions made one after another (an agent loop, a game tick, classify-then-act pipelines), latency budgets in milliseconds, boxes with no egress, or simply not wanting every classification to depend on one external API.
On Banking77 (77 customer-support intents), replayed against Jev's recorded answers, the local model answered 71.9% of held-out messages at 99.50% agreement with Jev (target 98%), taking over most traffic from about 5,000 messages on, and kept Jev's accuracy (78.5% vs Jev's 78.5% on the dataset's labels). The live run of 2026-09-25 gave 70.7% at 99.45%. A local answer takes ~15 ms on CPU (p50), about 20× faster than Jev; one CPU process answers ~130 messages/s, a GPU ~2,000/s.¹
Five public tasks, replayed through the loop with Jev's recorded answers as the teacher (bge-small on CPU, target agreement 98%). Answered locally and agreement are measured on 2,000 held-out rows against Jev; accuracy is against the dataset's own labels, for Jev alone and for the system (student where it answers, Jev elsewhere).
| Task | Answered locally | Agreement with Jev | Accuracy, Jev / system |
|---|---|---|---|
| Banking77, 77 intents | 71.9% | 99.50% | 78.5% / 78.5% |
| CLINC150, 150 intents † | 69.0% | 99.65% | 90.1% / 90.1% |
| AG News, 4 sections | 80.2% | 99.40% | 88.7% / 88.7% |
| TweetEval sentiment, 3 classes | 22.2% | 98.70% | 64.2% / 64.5% |
| TweetEval offensive, 2 classes | 24.1% | 98.80% | 73.8% / 74.2% |
† with rare_classes = "defer": Jev never used one of the task's labels.
Over 100 random splits across the 5 tasks, the calibrated threshold exceeded the 2% budget once; the usual point-estimate rule exceeded it on 6–12 of 20 splits per task.
Coverage tracks how consistent Jev itself is on a task, not how hard the task is: on the tweets Jev agrees with the dataset's labels only 64% and 74% of the time, so the student answers the confident quarter and forwards the rest, and the system's accuracy still matches Jev's. The threshold-rule comparison, what other targets buy, and the one-command reproduction: docs/benchmarks.md.
You set one number. Jevstiller returns the label Jev would have returned on at least that share of requests:
target_agreement = 0.98 → disagreement budget = 2% of all requests
The routing threshold is chosen on a held-out, IID calibration set so that, with 95% confidence, the share of requests the local model answers and gets different from Jev stays within the budget. It uses an exact finite-sample bound (Clopper–Pearson), testing candidate thresholds strictest-first. It is not tuned by eye, and it is re-verified forever on the audit channel. If the audit shows the contract is broken, everything falls back to Jev automatically.
If your code also acts on Jev's confidence, for example sending answers below 0.6 to review, set that number as the task's confidence_floor. The bound then also counts the requests Jev would have been less sure about than that, and local answers report at least that confidence. See configuration.
Agreement with Jev is not accuracy. If Jev is wrong, the student is wrong the same way. The status report says so next to every number. See DESIGN.md §2.
pip install jevstiller # everything: the proxy (`jevstiller serve`), the encoder, the Jev adapter
pip install "jevstiller[gpu]" # on a GPU machine: adds PyTorch, used automatically when CUDA is present
Python 3.10+. CPU works out of the box; a GPU only speeds up the encoder.
Run Jevstiller next to your services and point the Jev SDK at it:
docker run -d -p 8080:8080 -v jevstiller-data:/data ghcr.io/tomerglick57/jevstiller:0.4.0
# or: pip install jevstiller && jevstiller serve --config deploy/jevstiller.toml
export TYPESAFE_BASE_URL=http://jevstiller:8080 # in each calling service; nothing else changes
Services keep their own TYPESAFE_API_KEY; Jevstiller forwards it and never stores it.
Choice question becomes a task, keyed by its exact instructions, criteria and model, so services asking the same question share one local model.x-jevstiller-source: local tells you which), but only for keys Jev has accepted.Choice questions, other endpoints, and anything it does not understand.Performance of one process (16 vCPU, bge-small on CPU, docs/benchmarks.md):
Operations: a TOML config, an admin API and CLI (jevstiller admin ...), Prometheus /metrics, /readyz, JSON logs, jevstiller backup, and a Docker image and Kubernetes manifest. Docs: deploy, operations, the proxy, configuration, security, compatibility. Website: jevstiller.pages.dev.
As of September 2026:
stuntd train / stuntd enable.Without a key, python examples/quickstart_synthetic.py runs the whole loop on CPU in half a minute with a synthetic teacher standing in for Jev. With Jev (TYPESAFE_API_KEY in your environment):
from jevstiller import Task, Jevstiller, load_encoder
from jevstiller.teachers.jev import JevTeacher
task = Task(
name="support_router",
instructions="Which team should handle this customer message?",
classes={ # descriptions are sent to Jev verbatim — they are the spec
"billing": "Charges, invoices, refunds, payment methods",
"technical": "Bugs, errors, integrations, things not working",
"cancellation": "Wants to cancel, downgrade, or close the account",
"sales": "Pre-sales questions, plan comparison, quotes",
"other": "Anything that does not fit the categories above",
},
target_agreement=0.98,
)
js = Jevstiller(task, teacher=JevTeacher(), data_dir="./jevstiller-data", encoder=load_encoder("base"))
r = js.classify("Please cancel my subscription")
r.label # "cancellation"
r.confidence # 0.97
r.source # "teacher" at first — later "student:v9"
print(js.status().report())
Task: support_router version 89a7438c2f7d mode: cascade audit rate 2%
Production: student:v9 Shadow: -
Requests: 11,083 student 69.8% teacher 30.2%
Agreement with teacher (audit, n=415): 99.40% [98.46%, 99.84%] target 98% OK
note: agreement with the teacher is not accuracy.
bge-small/base/large), ONNX or PyTorch. Embeddings are stored, so retraining never re-encodes.student:vN directories; promote, roll back, or export() a standalone bundle.The full reasoning — including the five loop bugs the first real replay found and how they were fixed — is in DESIGN.md.
Alpha. Validated against live Jev (above).
Proxy: the drop-in proxy works end to end with the unmodified TypeSafe SDK.
Security: three security audits (12, 15 and 8 findings, all fixed: docs/security.md).
Deployment: it runs as a hardened container.
Tested under failure and load:
kill -9 during training.Results: docs/benchmarks.md.
Benchmarked on five public tasks against Jev's recorded answers (above); the full tables and the one-command reproduction are in docs/benchmarks.md.
See DEPLOYMENT_PLAN.md for what is done and what is next.
Documentation:
Not for: tasks with changing class lists (retrain from scratch), non-text input, or volumes too low to ever collect a few thousand examples.
Jev's answers for every Banking77 message ship with the repository, so the headline result replays without an API key:
git clone https://github.com/tomerglick57/Jevstiller && cd Jevstiller && pip install jevstiller
bash experiments/reproduce.sh # the Banking77 result above, on CPU, 10-15 minutes: 71.9% coverage at 99.50% agreement, deterministic
python experiments/run.py --dataset banking77 --teacher jev --encoder small # against live Jev (TYPESAFE_API_KEY)
python experiments/run.py --dataset banking77 --teacher oracle --encoder base # the dataset's labels as a perfect teacher
Apache 2.0 (see LICENSE and NOTICE). Jevstiller is an independent project and is not affiliated with, endorsed by, or supported by TypeSafe. "Jev" is their model; this tool only talks to its public API.
¹ Live run 2026-09-24: 11,083 replayed messages with 2,000 held out, bge-small on CPU; the GPU figure is from an oracle-teacher run with bge-base on an RTX 3090. Every number, with its command: docs/benchmarks.md.
Python
89.8%
Astro
3.5%
JavaScript
2.6%
CSS
2.0%
MDX
1.1%