pavanjava/jev_and_laya_benchmarking

Python

1

1 commits

updated Sep 25, 2026

See the code

README

Jev vs Laya on Typed Decisions

A benchmark of two System One decision models on LocalLLaMA/typed-decisions, measuring accuracy against gold and inference speed.

ModelHow it runs
JevTypeSafe System One API (typesafe-sdk), commercial GPUs, zero-shot
Layaconvaiinnovations/laya, typed-decisions weights, local Apple M1 Pro CPU

Dataset

  • 1,600 unique cases (1,200 train / 400 test) across 4 workflows: agent trace observability, customer service, invoice processing, security incidents.
  • Each case asks 5 typed questions about one shared state: choice (one of N labels), score (ordered rubric), noul (yes/no). That is 8,000 decisions per model.
  • Gold labels are the mean of 3 teacher-model samples. A prediction counts as correct when its most likely answer matches the gold label.
  • Hugging Face lists 3,200 rows because the all config duplicates the 4 per-workflow configs. The benchmark uses all only, so no case is counted twice.

Results (test split, 400 cases)

ModelAccuracychoicescorenoulp50 latencyp95 latency
Jev0.7340.7300.6990.783756 ms2,957 ms
Laya0.7660.7330.7230.857484 ms663 ms
  • Laya leads Jev by 3.2 points (95% bootstrap CI over cases: +1.1 to +5.5).
  • Jev's 0.734 is close to its published leaderboard score (0.727), and Laya's matches its model card (0.766), so the scoring is consistent with independent results.
  • Only the test split is a fair comparison. Laya's typed-decisions checkpoint was fine-tuned on this dataset's train split (per its model card), so it scores 0.854 on train vs 0.766 on test. Jev scores 0.734 on both.
  • These are different modes: Laya is fine-tuned on this benchmark, Jev answers without training on it. The dataset card says the two are not directly comparable.
  • Latency is not like-for-like: Laya runs locally on CPU, while Jev's time includes the network round trip.

Overall accuracy Accuracy by workflow Accuracy by question type Latency

Run it

Needs a TYPESAFE_API_KEY in .env.

UV_PROJECT_ENVIRONMENT=venv uv sync
venv/bin/python benchmark.py          # calls both models on all 1,600 cases
venv/bin/python results_analysis.py   # builds the charts and PDF from the CSV

The benchmark makes 1,600 sequential Jev API calls, so expect it to take 20+ minutes.

Files

FileWhat it is
benchmark.pyLoads the dataset, runs both models, writes the results CSVs
results_analysis.pyBuilds the charts and PDF report from results/benchmark.csv
results/benchmark.csvOne row per case, question and model: id, split, workflow, question, qtype, model, gold, pred, correct, latency_ms
results/summary.csvAggregated accuracy and latency
results/jev_vs_laya_report.pdfOne summary page plus the 4 charts
results/figures/*.pngThe 4 charts as images
jev_quick_start.py, laya_quick_start.pyMinimal usage examples for each model

License

MIT

Not covered

Accuracy only checks the most likely answer. Calibration metrics (Brier score, KL from the gold distribution, ECE) are not computed. Laya's model card reports that Jev matches the gold probability distributions better (0.580 vs 0.471), which this benchmark does not measure.

pavanjava/jev_and_laya_benchmarking

Python

1

1 commits

updated Sep 25, 2026

See the code

README

Jev vs Laya on Typed Decisions

A benchmark of two System One decision models on LocalLLaMA/typed-decisions, measuring accuracy against gold and inference speed.

ModelHow it runs
JevTypeSafe System One API (typesafe-sdk), commercial GPUs, zero-shot
Layaconvaiinnovations/laya, typed-decisions weights, local Apple M1 Pro CPU

Dataset

  • 1,600 unique cases (1,200 train / 400 test) across 4 workflows: agent trace observability, customer service, invoice processing, security incidents.
  • Each case asks 5 typed questions about one shared state: choice (one of N labels), score (ordered rubric), noul (yes/no). That is 8,000 decisions per model.
  • Gold labels are the mean of 3 teacher-model samples. A prediction counts as correct when its most likely answer matches the gold label.
  • Hugging Face lists 3,200 rows because the all config duplicates the 4 per-workflow configs. The benchmark uses all only, so no case is counted twice.

Results (test split, 400 cases)

ModelAccuracychoicescorenoulp50 latencyp95 latency
Jev0.7340.7300.6990.783756 ms2,957 ms
Laya0.7660.7330.7230.857484 ms663 ms
  • Laya leads Jev by 3.2 points (95% bootstrap CI over cases: +1.1 to +5.5).
  • Jev's 0.734 is close to its published leaderboard score (0.727), and Laya's matches its model card (0.766), so the scoring is consistent with independent results.
  • Only the test split is a fair comparison. Laya's typed-decisions checkpoint was fine-tuned on this dataset's train split (per its model card), so it scores 0.854 on train vs 0.766 on test. Jev scores 0.734 on both.
  • These are different modes: Laya is fine-tuned on this benchmark, Jev answers without training on it. The dataset card says the two are not directly comparable.
  • Latency is not like-for-like: Laya runs locally on CPU, while Jev's time includes the network round trip.

Overall accuracy Accuracy by workflow Accuracy by question type Latency

Run it

Needs a TYPESAFE_API_KEY in .env.

UV_PROJECT_ENVIRONMENT=venv uv sync
venv/bin/python benchmark.py          # calls both models on all 1,600 cases
venv/bin/python results_analysis.py   # builds the charts and PDF from the CSV

The benchmark makes 1,600 sequential Jev API calls, so expect it to take 20+ minutes.

Files

FileWhat it is
benchmark.pyLoads the dataset, runs both models, writes the results CSVs
results_analysis.pyBuilds the charts and PDF report from results/benchmark.csv
results/benchmark.csvOne row per case, question and model: id, split, workflow, question, qtype, model, gold, pred, correct, latency_ms
results/summary.csvAggregated accuracy and latency
results/jev_vs_laya_report.pdfOne summary page plus the 4 charts
results/figures/*.pngThe 4 charts as images
jev_quick_start.py, laya_quick_start.pyMinimal usage examples for each model

License

MIT

Not covered

Accuracy only checks the most likely answer. Calibration metrics (Brier score, KL from the gold distribution, ECE) are not computed. Laya's model card reports that Jev matches the gold probability distributions better (0.580 vs 0.471), which this benchmark does not measure.