System 0 local decision gateway powered by Gemma 4. Direct logit scoring, KV-cache sequence branching, and asymmetric audit trails via llama.cpp.
Python
3
3 commits
updated Sep 28, 2026
A high-speed, local decision gateway running on Gemma 4 (llama.cpp / CUDA) that replaces flaky autoregressive JSON generation with direct token logit scoring, mathematical cyclic debiasing, asymmetric post-decision quotation, and Platt temperature calibration.
Get deterministic classifications, well-calibrated confidence scores, and verbatim audit trails in 15–45ms fast-path classification on consumer hardware.
https://github.com/user-attachments/assets/a654e5ff-d425-4299-b5a1-9a3a1399d666
Autoregressive LLM classification pipelines suffer from four critical failure modes:
Standard Autoregressive Classification (500–2,500ms):
Prompt ──> Generated CoT / JSON ──> [Logit Poisoning Risk] ──> Brittle Output
Gevva0 Calibrated Dual-Path (15–45ms):
Prompt ──> Direct Logit Readout ──> Platt Scaled ──> Confidence >= Threshold?
│
├── YES ──> Fast-Path Locked Verdict (15-25ms)
└── NO ──> Bounded Verification (30-85ms) ─┘
│
[Verdict Permanently Locked]
│
▼ (Isolated Forward Pass)
Asymmetric Grounded Evidence Extraction
Evaluated across easy, original, and hard forensic legal batteries from the official JevBench v1.4.2 benchmark suite.
| Rank | System Architecture | JevBench Score | Intelligence | Hard Tier (Forensic) | p50 Latency |
|---|---|---|---|---|---|
| ⭐ #1 | Gevva0 (Gemma 4 26B-A4B MoE) | 74.63 | 87.2 | 82.9% | 214 ms |
| #2 | decider-4b v2 (Mapika) | 64.13 | 49.4 | 41.4% | 118 ms |
| #3 | TypeSafe Jev 1.13.0 (Closed API) | 63.29 | 53.1 | 47.7% | 652 ms |
| #4 | JevK5 v0.2.0 | 62.04 | 48.9 | 42.3% | 122 ms |
| #5 | Cygnet (Frozen Gemma-4-12B) | 61.76 | 49.5 | 43.2% | 126 ms |
See full evaluation report:
docs/JEVBENCH_PUBLICATION_EVALUATION_26B.md
All local conditions were evaluated under strictly isolated weights (gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf, 14 GB) across 650 balanced multi-class scenarios (including 18.5% out-of-distribution distractor controls).
| Metric | Baseline (Naive Logits) | Gevva0 (Fast-Path) | Gevva0 (Adaptive CoT) | TypeSafe Jev (Cloud API) |
|---|---|---|---|---|
| Top-1 Accuracy | 76.0% [72.8, 79.2] | 89.1% [86.6, 91.4] | 89.8% [87.5, 92.2] | 88.3% [85.8, 90.6] |
| ECE (10 Bins, Calibration) | 0.186 | 0.064 | 0.027 | 0.043 |
| Brier Score (Lower=Better) | 0.452 | 0.211 | 0.189 | 0.217 |
| Label Invariance (Debiased) | 71.5% | 100.0% | 100.0% | 90.3% |
| OOD False-Positive Rate | 22.5% | 5.0% | 3.3% | 8.3% |
| Latency (p50 / p95) | 18ms / 23ms | 22ms / 28ms | 45ms / 86ms | 112ms / 221ms |
| Significance (McNemar) | — | $p < 0.0001$ | $p < 0.0001$ | $p = 0.4152$ (par) |
See full evaluation report: Gemma 4 26B-A4B (MoE):
BENCHMARK_REPORT_gemma-4-26B-A4B.md
Requires Python 3.12 and an NVIDIA GPU (CUDA 12+ / 13+). Prebuilt wheels are bundled via llama-cpp-python:
# Using uv (Recommended)
uv sync
# Or using standard pip
pip install -r requirements.txt
Verify your GPU environment and model loading:
uv run gevva0 check
from gevva0 import GevvaEngine
# Initialize the engine (auto-discovers model from llm_config.json)
engine = GevvaEngine()
context = "Customer reports unauthorized double charge on order #89211 after payment timeout."
options = [
"Billing Dispute / Refund",
"Technical Bug / Gateway Timeout",
"Account Security Incident",
"General Inquiry"
]
# Run decision with cyclic debiasing and confidence calibration
result = engine.decide(
context=context,
options=options,
confidence_threshold=0.85,
cyclic_debias=True
)
print(f"Verdict: {result.selected_option}")
print(f"Confidence: {result.calibrated_p:.4f}")
print(f"Path Taken: {result.path}") # 'fast_path' or 'adaptive_cot'
# Asymmetric post-verdict quote extraction (zero logit poisoning)
audit = engine.extract_audit(context=context, locked_choice=result.selected_option)
print(f"Grounding Quote: '{audit.verbatim_quote}'")
# Fast-path decision outputting structured JSON
uv run gevva0 decide `
--context "Customer reports their invoice was charged twice for the same month." `
--options "A: Billing Inquiry,B: Technical Bug,C: Churn Risk" `
--json
# Force CoT scratchpad via high confidence gate threshold
uv run gevva0 decide --context "..." --options "A: x,B: y" --threshold 0.99
# Full cyclic label-permutation debiasing (invariance guarantee)
uv run gevva0 decide --context "..." --options "A: x,B: y,C: z" --cyclic
# Launch API service & Web Dashboard UI (http://localhost:8000/ui)
uv run gevva0 serve --host 127.0.0.1 --port 8000
llm.scores[llm.n_tokens - 1]. Marginalizes over bare and space-prefixed tokens (logsumexp(logit(" A"), logit("A"))) to capture true prior distributions without running an autoregressive decoding loop.For mathematical proofs, multi-class Brier score Murphy decompositions, and VRAM sizing charts, see
docs/ARCHITECTURE.md.
├── src/
│ └── gevva0/ # Core Python package
│ ├── engine.py # llama.cpp logit extraction, KV caching, calibration
│ ├── debias.py # Cyclic permutation debiasing
│ ├── audit.py # Asymmetric quote & rationale extractor
│ ├── calibration.py # Platt temperature calibration & Brier fitting
│ ├── metrics.py # Bootstrap CIs, McNemar tests, ECE, Brier decomposition
│ ├── config.py # Model path, context size, and auto-discovery
│ ├── schema.py # Pydantic request / response schemas
│ ├── server.py # FastAPI service (serves /ui and API routes)
│ └── cli.py # CLI entrypoint ('gevva0')
├── ui/
│ └── web_dashboard/ # Interactive dashboard & audit UI (mounted at / and /ui/)
│ └── index.html # Single source of truth web interface
├── benchmarks/
│ ├── benchmark_suite_650.json # Rigorous N=650 4-class balanced evaluation battery
│ ├── benchmark_suite.json # Standard benchmark suite
│ ├── generate_rigorous_dataset.py# Generator for N=650 dataset with 18.5% OOD controls
│ ├── run_benchmark.py # Standardized CLI runner with 4 ablation lines
│ ├── generate_fixtures.py # Generates multimodal test images
│ └── fixtures/ # Test image assets (AP invoices, 404 UI, CCTV frames)
├── tests/
│ ├── tests.json # Curated interactive showcase tests for Web UI
│ ├── test_metrics.py # Unit tests for statistical & calibration formulas
│ └── test_benchmark_runner.py # Integration tests for benchmark runner
├── docs/
│ ├── ARCHITECTURE.md # In-depth technical write-up on logit scoring & calibration
│ ├── BENCHMARK_REPORT_gemma-4-26B-A4B.md
│ └── JEVBENCH_PUBLICATION_EVALUATION_26B.md
├── models/
│ └── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/
│ ├── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
│ └── mmproj-BF16.gguf
├── requirements.txt # Autogenerated requirements for pip / non-uv environments
├── pyproject.toml # Project dependencies managed via uv
└── README.md
llm_config.json){
"n_ctx": 4096,
"n_batch": 2048,
"n_seq_max": 8,
"kv_unified": true,
"kv_cache": "F16",
"type_k": 1,
"type_v": 1,
"model": {
"model_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf",
"mmproj_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/mmproj-BF16.gguf",
"n_gpu_layers": -1,
"verbose": false
},
"decision": {
"cot_threshold": 0.85,
"cot_max_tokens": 256,
"cot_temp": 0.0,
"cot_prompt": "Analysis: First calculate everything: ",
"cyclic_debias": true,
"kv_branching": true,
"kv_branching_min_tokens": 50
}
}
Overrides can also be set via environment variables: GEVVA0_MODEL, GEVVA0_N_CTX, and GEVVA0_CALIBRATION.
The benchmark suite includes an automated runner with $B=10,000$ bootstrap resampling and McNemar continuity-corrected significance testing:
# Run benchmark on active model
uv run python benchmarks/run_benchmark.py
# Run on specific parameter scale
uv run python benchmarks/run_benchmark.py --model 26b # Gemma 4 26B-A4B MoE
uv run python benchmarks/run_benchmark.py --model e4b # Gemma 4 E4B Dense
uv run python benchmarks/run_benchmark.py --model e2b # Gemma 4 E2B Edge
# Re-generate synthetic N=650 balanced evaluation battery with OOD controls
uv run python benchmarks/generate_rigorous_dataset.py
Apache License 2.0. See LICENSE for details.
Python
77.1%
HTML
22.9%
System 0 local decision gateway powered by Gemma 4. Direct logit scoring, KV-cache sequence branching, and asymmetric audit trails via llama.cpp.
Python
3
3 commits
updated Sep 28, 2026
A high-speed, local decision gateway running on Gemma 4 (llama.cpp / CUDA) that replaces flaky autoregressive JSON generation with direct token logit scoring, mathematical cyclic debiasing, asymmetric post-decision quotation, and Platt temperature calibration.
Get deterministic classifications, well-calibrated confidence scores, and verbatim audit trails in 15–45ms fast-path classification on consumer hardware.
https://github.com/user-attachments/assets/a654e5ff-d425-4299-b5a1-9a3a1399d666
Autoregressive LLM classification pipelines suffer from four critical failure modes:
Standard Autoregressive Classification (500–2,500ms):
Prompt ──> Generated CoT / JSON ──> [Logit Poisoning Risk] ──> Brittle Output
Gevva0 Calibrated Dual-Path (15–45ms):
Prompt ──> Direct Logit Readout ──> Platt Scaled ──> Confidence >= Threshold?
│
├── YES ──> Fast-Path Locked Verdict (15-25ms)
└── NO ──> Bounded Verification (30-85ms) ─┘
│
[Verdict Permanently Locked]
│
▼ (Isolated Forward Pass)
Asymmetric Grounded Evidence Extraction
Evaluated across easy, original, and hard forensic legal batteries from the official JevBench v1.4.2 benchmark suite.
| Rank | System Architecture | JevBench Score | Intelligence | Hard Tier (Forensic) | p50 Latency |
|---|---|---|---|---|---|
| ⭐ #1 | Gevva0 (Gemma 4 26B-A4B MoE) | 74.63 | 87.2 | 82.9% | 214 ms |
| #2 | decider-4b v2 (Mapika) | 64.13 | 49.4 | 41.4% | 118 ms |
| #3 | TypeSafe Jev 1.13.0 (Closed API) | 63.29 | 53.1 | 47.7% | 652 ms |
| #4 | JevK5 v0.2.0 | 62.04 | 48.9 | 42.3% | 122 ms |
| #5 | Cygnet (Frozen Gemma-4-12B) | 61.76 | 49.5 | 43.2% | 126 ms |
See full evaluation report:
docs/JEVBENCH_PUBLICATION_EVALUATION_26B.md
All local conditions were evaluated under strictly isolated weights (gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf, 14 GB) across 650 balanced multi-class scenarios (including 18.5% out-of-distribution distractor controls).
| Metric | Baseline (Naive Logits) | Gevva0 (Fast-Path) | Gevva0 (Adaptive CoT) | TypeSafe Jev (Cloud API) |
|---|---|---|---|---|
| Top-1 Accuracy | 76.0% [72.8, 79.2] | 89.1% [86.6, 91.4] | 89.8% [87.5, 92.2] | 88.3% [85.8, 90.6] |
| ECE (10 Bins, Calibration) | 0.186 | 0.064 | 0.027 | 0.043 |
| Brier Score (Lower=Better) | 0.452 | 0.211 | 0.189 | 0.217 |
| Label Invariance (Debiased) | 71.5% | 100.0% | 100.0% | 90.3% |
| OOD False-Positive Rate | 22.5% | 5.0% | 3.3% | 8.3% |
| Latency (p50 / p95) | 18ms / 23ms | 22ms / 28ms | 45ms / 86ms | 112ms / 221ms |
| Significance (McNemar) | — | $p < 0.0001$ | $p < 0.0001$ | $p = 0.4152$ (par) |
See full evaluation report: Gemma 4 26B-A4B (MoE):
BENCHMARK_REPORT_gemma-4-26B-A4B.md
Requires Python 3.12 and an NVIDIA GPU (CUDA 12+ / 13+). Prebuilt wheels are bundled via llama-cpp-python:
# Using uv (Recommended)
uv sync
# Or using standard pip
pip install -r requirements.txt
Verify your GPU environment and model loading:
uv run gevva0 check
from gevva0 import GevvaEngine
# Initialize the engine (auto-discovers model from llm_config.json)
engine = GevvaEngine()
context = "Customer reports unauthorized double charge on order #89211 after payment timeout."
options = [
"Billing Dispute / Refund",
"Technical Bug / Gateway Timeout",
"Account Security Incident",
"General Inquiry"
]
# Run decision with cyclic debiasing and confidence calibration
result = engine.decide(
context=context,
options=options,
confidence_threshold=0.85,
cyclic_debias=True
)
print(f"Verdict: {result.selected_option}")
print(f"Confidence: {result.calibrated_p:.4f}")
print(f"Path Taken: {result.path}") # 'fast_path' or 'adaptive_cot'
# Asymmetric post-verdict quote extraction (zero logit poisoning)
audit = engine.extract_audit(context=context, locked_choice=result.selected_option)
print(f"Grounding Quote: '{audit.verbatim_quote}'")
# Fast-path decision outputting structured JSON
uv run gevva0 decide `
--context "Customer reports their invoice was charged twice for the same month." `
--options "A: Billing Inquiry,B: Technical Bug,C: Churn Risk" `
--json
# Force CoT scratchpad via high confidence gate threshold
uv run gevva0 decide --context "..." --options "A: x,B: y" --threshold 0.99
# Full cyclic label-permutation debiasing (invariance guarantee)
uv run gevva0 decide --context "..." --options "A: x,B: y,C: z" --cyclic
# Launch API service & Web Dashboard UI (http://localhost:8000/ui)
uv run gevva0 serve --host 127.0.0.1 --port 8000
llm.scores[llm.n_tokens - 1]. Marginalizes over bare and space-prefixed tokens (logsumexp(logit(" A"), logit("A"))) to capture true prior distributions without running an autoregressive decoding loop.For mathematical proofs, multi-class Brier score Murphy decompositions, and VRAM sizing charts, see
docs/ARCHITECTURE.md.
├── src/
│ └── gevva0/ # Core Python package
│ ├── engine.py # llama.cpp logit extraction, KV caching, calibration
│ ├── debias.py # Cyclic permutation debiasing
│ ├── audit.py # Asymmetric quote & rationale extractor
│ ├── calibration.py # Platt temperature calibration & Brier fitting
│ ├── metrics.py # Bootstrap CIs, McNemar tests, ECE, Brier decomposition
│ ├── config.py # Model path, context size, and auto-discovery
│ ├── schema.py # Pydantic request / response schemas
│ ├── server.py # FastAPI service (serves /ui and API routes)
│ └── cli.py # CLI entrypoint ('gevva0')
├── ui/
│ └── web_dashboard/ # Interactive dashboard & audit UI (mounted at / and /ui/)
│ └── index.html # Single source of truth web interface
├── benchmarks/
│ ├── benchmark_suite_650.json # Rigorous N=650 4-class balanced evaluation battery
│ ├── benchmark_suite.json # Standard benchmark suite
│ ├── generate_rigorous_dataset.py# Generator for N=650 dataset with 18.5% OOD controls
│ ├── run_benchmark.py # Standardized CLI runner with 4 ablation lines
│ ├── generate_fixtures.py # Generates multimodal test images
│ └── fixtures/ # Test image assets (AP invoices, 404 UI, CCTV frames)
├── tests/
│ ├── tests.json # Curated interactive showcase tests for Web UI
│ ├── test_metrics.py # Unit tests for statistical & calibration formulas
│ └── test_benchmark_runner.py # Integration tests for benchmark runner
├── docs/
│ ├── ARCHITECTURE.md # In-depth technical write-up on logit scoring & calibration
│ ├── BENCHMARK_REPORT_gemma-4-26B-A4B.md
│ └── JEVBENCH_PUBLICATION_EVALUATION_26B.md
├── models/
│ └── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/
│ ├── gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf
│ └── mmproj-BF16.gguf
├── requirements.txt # Autogenerated requirements for pip / non-uv environments
├── pyproject.toml # Project dependencies managed via uv
└── README.md
llm_config.json){
"n_ctx": 4096,
"n_batch": 2048,
"n_seq_max": 8,
"kv_unified": true,
"kv_cache": "F16",
"type_k": 1,
"type_v": 1,
"model": {
"model_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf",
"mmproj_path": "models/gemma-4-26B-A4B-it-qat-UD-Q4_K_XL/mmproj-BF16.gguf",
"n_gpu_layers": -1,
"verbose": false
},
"decision": {
"cot_threshold": 0.85,
"cot_max_tokens": 256,
"cot_temp": 0.0,
"cot_prompt": "Analysis: First calculate everything: ",
"cyclic_debias": true,
"kv_branching": true,
"kv_branching_min_tokens": 50
}
}
Overrides can also be set via environment variables: GEVVA0_MODEL, GEVVA0_N_CTX, and GEVVA0_CALIBRATION.
The benchmark suite includes an automated runner with $B=10,000$ bootstrap resampling and McNemar continuity-corrected significance testing:
# Run benchmark on active model
uv run python benchmarks/run_benchmark.py
# Run on specific parameter scale
uv run python benchmarks/run_benchmark.py --model 26b # Gemma 4 26B-A4B MoE
uv run python benchmarks/run_benchmark.py --model e4b # Gemma 4 E4B Dense
uv run python benchmarks/run_benchmark.py --model e2b # Gemma 4 E2B Edge
# Re-generate synthetic N=650 balanced evaluation battery with OOD controls
uv run python benchmarks/generate_rigorous_dataset.py
Apache License 2.0. See LICENSE for details.
Python
77.1%
HTML
22.9%