yuhai-china/JEV-27B-DEMO

JEV-27B demos: one engine, two systems. Search re-ranking, agent decisions, System 1 to System 2, response judge, hallucination guard, zero-shot news and image recommendation.

Python

1

22 commits

updated Oct 2, 2026

See the code

README

JEV-27B demos: one engine, two systems

autotrust/JEV-27B serves two ways of answering from one set of weights in one vLLM engine:

what it doesoutputtypical time
System 1typed decisions: yes/no · pick one of 2-16 options · rate 0-5a calibrated probability for every option, in one forward pass~0.1 s
System 2the unmodified Qwen3.8-27B, optionally thinking step by steptext / reasoningseconds

This repository shows what that is good for, with real data and measured results.

中文亮点说明:HIGHLIGHTS_zh.md

#demoheadline result
01Search re-rankingnDCG@10 on TREC-COVID 0.858 vs 0.793 for bge-reranker-v2-m3 and 0.623 for BM25
02Agent decisions: triage, phishing, moderation, tool routing24 decisions in 0.43 s, no output parsing
03System 1 → System 2 escalation70% of questions answered in 0.1 s; accuracy 0.792 → 0.892 (thinking on everything: 0.917)
04Response judge / reward modelRewardBench 89.9 in one forward pass, ahead of GPT-4o (86.7), Gemini 1.5 Pro (88.2), Claude 3.5 Sonnet (84.2) as judges
05Hallucination guardanswer only the half System 1 trusts: accuracy 71% → 96% (System 2's own confidence: 87%)
06News recommendation, zero-shotnever trained on MIND or click data, AUC 0.642: beats every zero-shot baseline and LightGBM rankers trained on MIND
07Short-video recommendation, zero-shot (TikTok-style feeds)looks only at video covers, AUC 0.727: equals collaborative filtering learned from 59,045 users' logs
08Biomedical research questions (PubMedQA)77.8%, level with human experts (78.0%), zero-shot in one forward pass; above GPT-4's zero-shot 75.2%
09Multimodal judge78.3% on VL-RewardBench, above every model on its 2025 leaderboard; on the 2026 MMRB2, text-to-image 69.2 (GPT-5 70.5), average at GPT-4.1 level
10Agent judgePlan-RewardBench (ACL 2026): 73.2%, top of the table, ahead of GPT-5 (68.5) and Gemini-3-Flash (69.1); AgentRewardBench: higher precision than every leaderboard judge at its recall
appWeb app (Gradio)all of the above, interactive

search

system1 to system2

judge

guard

news

images

pubmedqa

multimodal judge

MMRB2

Plan-RewardBench

agent judge

Quick start

You need one GPU with 80 GB or more (H100 / H200 / B200 / RTX PRO 6000) for the server. The demos themselves run anywhere.

pip install -r requirements.txt            # demo side
pip install vllm                           # server side (Qwen3.5/3.8 support required)

bash common/serve_jev27b.sh                # downloads autotrust/JEV-27B (~54 GB) and starts vLLM on :8000
python 01-search-ranking/demo.py           # any demo
python app/app.py                          # web app on http://localhost:7860

Set JEV_URL if the server is not on localhost:8000.

Image input (demo 07): autotrust/JEV-27B-VL is JEV-27B with vision. bash common/serve_jev27b_mm.sh serves it, so both systems also accept images. It runs every other demo too.

System 1 as a plain HTTP endpoint: serve_jev27b_mm.sh starts vLLM through common/serve_decide.py, the standard vLLM OpenAI server plus POST /v1/decide. Send the question, get calibrated probabilities back; the server builds the prompt and applies the decision head. Same request and response format as the hosted API, with 2–256 options per choice question and images allowed in state.

curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
  "kind": "choice", "state": "Customer: my card was charged twice for one coffee.",
  "question": "Which team should handle this?", "options": ["billing", "shipping", "tech support"]}'
# {"options": [...], "probabilities": [0.9979, 0.00001, 0.0021], "choice": "billing", "choice_index": 0, ...}

kind is noul (yes/no), score (0–5) or choice. For images, make state a list: ["Picture: ", {"image": "https://... or data:image/png;base64,..."}]. JEV_BACKEND=decide makes jev_client use this endpoint too. Over 16 options, the labels continue past A–P (Q–Z, AA, AB, …). On CLINC150 with all 150 intents as options, accuracy is 89.5% with intent names alone and 93.8% when each option gets a one-line description.

Using a hosted JEV API instead of your own GPU: set the URL and key, then run any demo or the app unchanged.

export JEV_URL="https://jev-h200.scienceguru.ai/v1"
export JEV_API_KEY="<your API key>"      # System 1 then uses POST /v1/decide; all requests send the Bearer key

Setting JEV_API_KEY selects the hosted backend automatically; JEV_BACKEND=vllm|decide overrides it. Details (Chinese): HIGHLIGHTS_zh.md → 切换 API 服务.

The whole API

from jev_client import decide, decide_many, chat, split_thinking      # common/jev_client.py, ~80 lines

decide("choice", state, "Which team should handle this ticket?", ["billing", "mobile app", "platform / SSO"])
# e.g. {'billing': 0.998, 'mobile app': 0.001, 'platform / SSO': 0.001}
decide("noul", state, "Is this scenario one where: a human must respond personally?")    # -> {'false': .., 'true': ..}
decide("score", state, "Rate how urgent this ticket is on a 0-5 scale.")                 # -> {'0': .., ..., '5': ..}
decide_many([(kind, state, question, options), ...])      # concurrent; vLLM batches them on the GPU
decide_mm("noul", ["Covers the user watched:", {"image": "a.jpg"}, "Candidate:", {"image": "b.jpg"}],
          "Is this scenario one where: this user clicks on the candidate video?")        # images (multimodal server)

reasoning, answer = split_thinking(chat(prompt, thinking=True))                          # System 2

state is any text or JSON. System 1 is a LoRA adapter plus a decision head expressed as an lm_head LoRA, served by vLLM as the model jev-decision. jev_client.decide asks for the logprobs of the verbalizer tokens only and applies the bundled bias and per-kind temperature, so no text is ever generated or parsed.

Screenshots

playgroundsearch
escalationphishing
judgeguard
newsimages
biomedical

Licence

Code: Apache-2.0. Model: see autotrust/JEV-27B. Data: TREC-COVID / NFCorpus via BEIR, GSM8K (MIT), AQuA-RAT (Apache-2.0), ARC (CC BY-SA 4.0), CommonsenseQA (MIT), RewardBench (ODC-BY), TriviaQA (Apache-2.0), PubMedQA (MIT), VL-RewardBench (research use, downloaded at run time), AgentRewardBench and Plan-RewardBench (CC BY 4.0) (downloaded at run time), MIND (Microsoft Research License Terms) and MicroLens (Westlake University, research use), both downloaded at run time and not redistributed.

yuhai-china/JEV-27B-DEMO

JEV-27B demos: one engine, two systems. Search re-ranking, agent decisions, System 1 to System 2, response judge, hallucination guard, zero-shot news and image recommendation.

Python

1

22 commits

updated Oct 2, 2026

See the code

README

JEV-27B demos: one engine, two systems

autotrust/JEV-27B serves two ways of answering from one set of weights in one vLLM engine:

what it doesoutputtypical time
System 1typed decisions: yes/no · pick one of 2-16 options · rate 0-5a calibrated probability for every option, in one forward pass~0.1 s
System 2the unmodified Qwen3.8-27B, optionally thinking step by steptext / reasoningseconds

This repository shows what that is good for, with real data and measured results.

中文亮点说明:HIGHLIGHTS_zh.md

#demoheadline result
01Search re-rankingnDCG@10 on TREC-COVID 0.858 vs 0.793 for bge-reranker-v2-m3 and 0.623 for BM25
02Agent decisions: triage, phishing, moderation, tool routing24 decisions in 0.43 s, no output parsing
03System 1 → System 2 escalation70% of questions answered in 0.1 s; accuracy 0.792 → 0.892 (thinking on everything: 0.917)
04Response judge / reward modelRewardBench 89.9 in one forward pass, ahead of GPT-4o (86.7), Gemini 1.5 Pro (88.2), Claude 3.5 Sonnet (84.2) as judges
05Hallucination guardanswer only the half System 1 trusts: accuracy 71% → 96% (System 2's own confidence: 87%)
06News recommendation, zero-shotnever trained on MIND or click data, AUC 0.642: beats every zero-shot baseline and LightGBM rankers trained on MIND
07Short-video recommendation, zero-shot (TikTok-style feeds)looks only at video covers, AUC 0.727: equals collaborative filtering learned from 59,045 users' logs
08Biomedical research questions (PubMedQA)77.8%, level with human experts (78.0%), zero-shot in one forward pass; above GPT-4's zero-shot 75.2%
09Multimodal judge78.3% on VL-RewardBench, above every model on its 2025 leaderboard; on the 2026 MMRB2, text-to-image 69.2 (GPT-5 70.5), average at GPT-4.1 level
10Agent judgePlan-RewardBench (ACL 2026): 73.2%, top of the table, ahead of GPT-5 (68.5) and Gemini-3-Flash (69.1); AgentRewardBench: higher precision than every leaderboard judge at its recall
appWeb app (Gradio)all of the above, interactive

search

system1 to system2

judge

guard

news

images

pubmedqa

multimodal judge

MMRB2

Plan-RewardBench

agent judge

Quick start

You need one GPU with 80 GB or more (H100 / H200 / B200 / RTX PRO 6000) for the server. The demos themselves run anywhere.

pip install -r requirements.txt            # demo side
pip install vllm                           # server side (Qwen3.5/3.8 support required)

bash common/serve_jev27b.sh                # downloads autotrust/JEV-27B (~54 GB) and starts vLLM on :8000
python 01-search-ranking/demo.py           # any demo
python app/app.py                          # web app on http://localhost:7860

Set JEV_URL if the server is not on localhost:8000.

Image input (demo 07): autotrust/JEV-27B-VL is JEV-27B with vision. bash common/serve_jev27b_mm.sh serves it, so both systems also accept images. It runs every other demo too.

System 1 as a plain HTTP endpoint: serve_jev27b_mm.sh starts vLLM through common/serve_decide.py, the standard vLLM OpenAI server plus POST /v1/decide. Send the question, get calibrated probabilities back; the server builds the prompt and applies the decision head. Same request and response format as the hosted API, with 2–256 options per choice question and images allowed in state.

curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
  "kind": "choice", "state": "Customer: my card was charged twice for one coffee.",
  "question": "Which team should handle this?", "options": ["billing", "shipping", "tech support"]}'
# {"options": [...], "probabilities": [0.9979, 0.00001, 0.0021], "choice": "billing", "choice_index": 0, ...}

kind is noul (yes/no), score (0–5) or choice. For images, make state a list: ["Picture: ", {"image": "https://... or data:image/png;base64,..."}]. JEV_BACKEND=decide makes jev_client use this endpoint too. Over 16 options, the labels continue past A–P (Q–Z, AA, AB, …). On CLINC150 with all 150 intents as options, accuracy is 89.5% with intent names alone and 93.8% when each option gets a one-line description.

Using a hosted JEV API instead of your own GPU: set the URL and key, then run any demo or the app unchanged.

export JEV_URL="https://jev-h200.scienceguru.ai/v1"
export JEV_API_KEY="<your API key>"      # System 1 then uses POST /v1/decide; all requests send the Bearer key

Setting JEV_API_KEY selects the hosted backend automatically; JEV_BACKEND=vllm|decide overrides it. Details (Chinese): HIGHLIGHTS_zh.md → 切换 API 服务.

The whole API

from jev_client import decide, decide_many, chat, split_thinking      # common/jev_client.py, ~80 lines

decide("choice", state, "Which team should handle this ticket?", ["billing", "mobile app", "platform / SSO"])
# e.g. {'billing': 0.998, 'mobile app': 0.001, 'platform / SSO': 0.001}
decide("noul", state, "Is this scenario one where: a human must respond personally?")    # -> {'false': .., 'true': ..}
decide("score", state, "Rate how urgent this ticket is on a 0-5 scale.")                 # -> {'0': .., ..., '5': ..}
decide_many([(kind, state, question, options), ...])      # concurrent; vLLM batches them on the GPU
decide_mm("noul", ["Covers the user watched:", {"image": "a.jpg"}, "Candidate:", {"image": "b.jpg"}],
          "Is this scenario one where: this user clicks on the candidate video?")        # images (multimodal server)

reasoning, answer = split_thinking(chat(prompt, thinking=True))                          # System 2

state is any text or JSON. System 1 is a LoRA adapter plus a decision head expressed as an lm_head LoRA, served by vLLM as the model jev-decision. jev_client.decide asks for the logprobs of the verbalizer tokens only and applies the bundled bias and per-kind temperature, so no text is ever generated or parsed.

Screenshots

playgroundsearch
escalationphishing
judgeguard
newsimages
biomedical

Licence

Code: Apache-2.0. Model: see autotrust/JEV-27B. Data: TREC-COVID / NFCorpus via BEIR, GSM8K (MIT), AQuA-RAT (Apache-2.0), ARC (CC BY-SA 4.0), CommonsenseQA (MIT), RewardBench (ODC-BY), TriviaQA (Apache-2.0), PubMedQA (MIT), VL-RewardBench (research use, downloaded at run time), AgentRewardBench and Plan-RewardBench (CC BY 4.0) (downloaded at run time), MIND (Microsoft Research License Terms) and MicroLens (Westlake University, research use), both downloaded at run time and not redistributed.