solvi-ai/solvi-base

Model

solvi-ai/solvi-base (preview)

0

3 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

solvi-ai/solvi-base (preview)

A 150M-parameter ModernBERT-base cross-encoder for solvi typed questions (same format and answer kinds as solvi-large), distilled from solvi-large for CPU / browser use: 50 ms per question on a CPU (ONNX fp16).

Honest summary. On typed questions over JSON states it matches the large model (typed-decisions 54.5%), and its contract evidence is much better than the previous base model (46% vs 23% supporting quotes). On zero-shot choice questions it is not better than the previous base models (Fast Decisions dev 56.3%; L14d 57.3%, GLiNER2.5-Decide 62.9%).

Results (held-out tests; never used for training or checkpoint selection)

Fast Decisions numbers are on the public dev split in our harness (Fastino's official numbers are on a hidden test split).

testsolvi-base (this)solvi-large (teacher)L14g (previous base)
Fast Decisions dev, zero-shot / Zc / Sc k=3256.3% / 58.2% / 60.2%59.4% / 60.5% / 62.0%56.2% / 57.4% / 58.8%
typed-decisions, zero-shot / k=30054.5% / 64.0%54.5% / 64.9%45.1% / 61.3%
ContractNLI (in-distribution), three answers87.3%88.8%56.0%
ContractNLI, evidence supports the answer46.4%62.0%23.3%
Taskmaster-2 held-out slots, span F1 / exact0.77 / 66.1%0.83 / 70.2%0.80 / 67.7%
synthetic held-out schemas, all kinds94.9%97.9%94.4%
act AUROC: Fast Decisions / typed-decisions / ContractNLI / Taskmaster-20.74 / 0.60 / 0.86 / 0.640.76 / 0.67 / 0.85 / 0.740.71 / 0.57 / 0.63 / 0.64
ECE: typed-decisions / ContractNLI / Taskmaster-20.12 / 0.02 / 0.110.21 / 0.03 / 0.060.19 / 0.04 / 0.10

For reference: GLiNER2.5-Decide 62.9% on Fast Decisions dev and 50.3% on typed-decisions; Laya base 36.0% on typed-decisions.

Speed (CPU, 4 threads): ONNX fp16, one question with 10 options and ~20 words: 50 ms (fp32 46 ms); a question over a long JSON state ~170–270 ms. Block layout (several questions per pass) agrees with one question per pass only 90.2% on typed-decisions, so multi_question.enabled = false; the block ONNX export is included for experiments.

Training

  • Initial weights: L14g (our previous base, ModernBERT-base lineage, Apache-2.0).
  • Distillation: 65 minutes on one GPU (≈ 9.3k updates, batch 64) on the same mix as solvi-large; every target is 0.5 × the data label + 0.5 × solvi-large's probabilities. Checkpoint chosen on the same development metric as the large model, never on the tests.
  • Data and teachers: as solvi-large (see its card): votes of Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501 and GLiNER2.5-Decide (all Apache-2.0), a Qwen2.5-32B data agent verified by two other models, ContractNLI train, VitaminC, SQuAD 2.0, BoolQ, synthetic typed states; de-duplicated against every test set.

Training data and attribution

The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA, with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.

sourcelicenseattribution
solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text)Apache-2.0solvi authors
SQuAD 2.0 (train)CC BY-SA 4.0Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018
BoolQ (train)CC BY-SA 3.0Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019
CFPB Consumer Complaint DatabaseCC0-1.0Consumer Financial Protection Bureau
English Wikinews (2023-07-28)CC BY 2.5Wikinews contributors, https://en.wikinews.org
arXiv abstracts (metadata)CC0-1.0arXiv.org
Taskmaster-1CC BY 4.0Google Research
MultiWOZ 2.2MITBudzianowski et al.
Amazon polarityApache-2.0Zhang et al.; McAuley & Leskovec
GoEmotionsApache-2.0Google Research
LEDGAR, UNFAIR-ToS (LexGLUE)CC BY 4.0Tuggener et al.; Lippi et al.; Chalkidis et al.
SMS Spam CollectionCC BY 4.0Almeida & Hidalgo (UCI)
Toxic conversations 50k (Civil Comments / Jigsaw)CC BY 4.0Jigsaw / Conversation AI; MTEB
deepset prompt-injectionsApache-2.0deepset
CLINC150CC BY 3.0Larson et al.
Banking77CC BY 4.0PolyAI
MASSIVE (en-US)CC BY 4.0Amazon Science
SNIPS NLUApache-2.0Snips
Bitext customer-support, retail-banking, travel, insuranceCDLA-Sharing-1.0Bitext Innovations

| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |

SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.

Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).

Independent benchmarks (as released, zero-shot)

benchmarksolvi-largesolvi-baseothers (their published numbers or the benchmark's leaderboard)
jabr classifier-benchmark v2 (49 tasks, 866 cases), macro0.6940.602Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583
decision-models-under-pressure: accuracy with 128 options45.1%50.8%Jev 60.0%, best other open model 41.0%, Laya 38.5%
same: answers changed by reordering 64 options40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default)24.1% as listedJev 14.6%, Laya 49.4%
same: accuracy lost to near-duplicate distractors (64 options)0.2600.247Jev 0.105, Laya 0.345

Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold labels have 50% — calibrate the threshold on your data (act_guard, fit).

Escalation with a guarantee (solvi ≥ 0.5.1)

The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real text. Calibrate on a few hundred labelled examples of your own stream instead:

part.act_guard(examples, risk=0.10)   # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples

Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:

data setshipped "10%" threshold: answered / erroract_guard(risk=0.10): answered / error / riskact_guard(risk=0.05)
typed-decisions66% / 40%27% / 36% / 9.9%16% / 30% / 5.0%
Taskmaster-284% / 37%39% / 25% / 10.1%25% / 20% / 5.0%
ContractNLI95% / 11%93% / 10% / 9.6%80% / 6% / 4.8%
JSON, 9 held-out schemas97% / 3%99.7% / 4.8% / 4.8%99.0% / 4.2% / 4.2%

The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.

Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.

Limitations

  • Zero-shot choice questions are not improved over the previous base models (Fast Decisions dev 56.3%; GLiNER2.5-Decide 62.9%). Fit on 30–60 examples of your task (Sc k=32: 60.2%).
  • Contract evidence supports the answer less often than the large model (46% vs 62%); act AUROC on typed-decisions 0.60.
  • The act / escalate thresholds shipped with the model are indicative only: fitted on development data, they are over-confident on new real text (see solvi-large). For a guaranteed risk, calibrate on your own labelled stream with act_guard (solvi ≥ 0.5.1, table above).
  • Several questions per pass disagree with one-question-per-pass in ~10% of answers on real states — disabled by default.
  • English only; truncation beyond 512 tokens (1024 in block layout). Not for medical, legal or credit decisions on its own.

Files

model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]), onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask), onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures, act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format), LICENSES.json (training data manifest).

Citation

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
calibration
classification
decision
distillation
extractive-qa
modernbert
onnx
safetensors
solvi
text-classification
typed-questions

solvi-ai/solvi-base

Model

solvi-ai/solvi-base (preview)

0

3 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

solvi-ai/solvi-base (preview)

A 150M-parameter ModernBERT-base cross-encoder for solvi typed questions (same format and answer kinds as solvi-large), distilled from solvi-large for CPU / browser use: 50 ms per question on a CPU (ONNX fp16).

Honest summary. On typed questions over JSON states it matches the large model (typed-decisions 54.5%), and its contract evidence is much better than the previous base model (46% vs 23% supporting quotes). On zero-shot choice questions it is not better than the previous base models (Fast Decisions dev 56.3%; L14d 57.3%, GLiNER2.5-Decide 62.9%).

Results (held-out tests; never used for training or checkpoint selection)

Fast Decisions numbers are on the public dev split in our harness (Fastino's official numbers are on a hidden test split).

testsolvi-base (this)solvi-large (teacher)L14g (previous base)
Fast Decisions dev, zero-shot / Zc / Sc k=3256.3% / 58.2% / 60.2%59.4% / 60.5% / 62.0%56.2% / 57.4% / 58.8%
typed-decisions, zero-shot / k=30054.5% / 64.0%54.5% / 64.9%45.1% / 61.3%
ContractNLI (in-distribution), three answers87.3%88.8%56.0%
ContractNLI, evidence supports the answer46.4%62.0%23.3%
Taskmaster-2 held-out slots, span F1 / exact0.77 / 66.1%0.83 / 70.2%0.80 / 67.7%
synthetic held-out schemas, all kinds94.9%97.9%94.4%
act AUROC: Fast Decisions / typed-decisions / ContractNLI / Taskmaster-20.74 / 0.60 / 0.86 / 0.640.76 / 0.67 / 0.85 / 0.740.71 / 0.57 / 0.63 / 0.64
ECE: typed-decisions / ContractNLI / Taskmaster-20.12 / 0.02 / 0.110.21 / 0.03 / 0.060.19 / 0.04 / 0.10

For reference: GLiNER2.5-Decide 62.9% on Fast Decisions dev and 50.3% on typed-decisions; Laya base 36.0% on typed-decisions.

Speed (CPU, 4 threads): ONNX fp16, one question with 10 options and ~20 words: 50 ms (fp32 46 ms); a question over a long JSON state ~170–270 ms. Block layout (several questions per pass) agrees with one question per pass only 90.2% on typed-decisions, so multi_question.enabled = false; the block ONNX export is included for experiments.

Training

  • Initial weights: L14g (our previous base, ModernBERT-base lineage, Apache-2.0).
  • Distillation: 65 minutes on one GPU (≈ 9.3k updates, batch 64) on the same mix as solvi-large; every target is 0.5 × the data label + 0.5 × solvi-large's probabilities. Checkpoint chosen on the same development metric as the large model, never on the tests.
  • Data and teachers: as solvi-large (see its card): votes of Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501 and GLiNER2.5-Decide (all Apache-2.0), a Qwen2.5-32B data agent verified by two other models, ContractNLI train, VitaminC, SQuAD 2.0, BoolQ, synthetic typed states; de-duplicated against every test set.

Training data and attribution

The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA, with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.

sourcelicenseattribution
solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text)Apache-2.0solvi authors
SQuAD 2.0 (train)CC BY-SA 4.0Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018
BoolQ (train)CC BY-SA 3.0Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019
CFPB Consumer Complaint DatabaseCC0-1.0Consumer Financial Protection Bureau
English Wikinews (2023-07-28)CC BY 2.5Wikinews contributors, https://en.wikinews.org
arXiv abstracts (metadata)CC0-1.0arXiv.org
Taskmaster-1CC BY 4.0Google Research
MultiWOZ 2.2MITBudzianowski et al.
Amazon polarityApache-2.0Zhang et al.; McAuley & Leskovec
GoEmotionsApache-2.0Google Research
LEDGAR, UNFAIR-ToS (LexGLUE)CC BY 4.0Tuggener et al.; Lippi et al.; Chalkidis et al.
SMS Spam CollectionCC BY 4.0Almeida & Hidalgo (UCI)
Toxic conversations 50k (Civil Comments / Jigsaw)CC BY 4.0Jigsaw / Conversation AI; MTEB
deepset prompt-injectionsApache-2.0deepset
CLINC150CC BY 3.0Larson et al.
Banking77CC BY 4.0PolyAI
MASSIVE (en-US)CC BY 4.0Amazon Science
SNIPS NLUApache-2.0Snips
Bitext customer-support, retail-banking, travel, insuranceCDLA-Sharing-1.0Bitext Innovations

| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |

SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.

Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).

Independent benchmarks (as released, zero-shot)

benchmarksolvi-largesolvi-baseothers (their published numbers or the benchmark's leaderboard)
jabr classifier-benchmark v2 (49 tasks, 866 cases), macro0.6940.602Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583
decision-models-under-pressure: accuracy with 128 options45.1%50.8%Jev 60.0%, best other open model 41.0%, Laya 38.5%
same: answers changed by reordering 64 options40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default)24.1% as listedJev 14.6%, Laya 49.4%
same: accuracy lost to near-duplicate distractors (64 options)0.2600.247Jev 0.105, Laya 0.345

Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold labels have 50% — calibrate the threshold on your data (act_guard, fit).

Escalation with a guarantee (solvi ≥ 0.5.1)

The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real text. Calibrate on a few hundred labelled examples of your own stream instead:

part.act_guard(examples, risk=0.10)   # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples

Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:

data setshipped "10%" threshold: answered / erroract_guard(risk=0.10): answered / error / riskact_guard(risk=0.05)
typed-decisions66% / 40%27% / 36% / 9.9%16% / 30% / 5.0%
Taskmaster-284% / 37%39% / 25% / 10.1%25% / 20% / 5.0%
ContractNLI95% / 11%93% / 10% / 9.6%80% / 6% / 4.8%
JSON, 9 held-out schemas97% / 3%99.7% / 4.8% / 4.8%99.0% / 4.2% / 4.2%

The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.

Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.

Limitations

  • Zero-shot choice questions are not improved over the previous base models (Fast Decisions dev 56.3%; GLiNER2.5-Decide 62.9%). Fit on 30–60 examples of your task (Sc k=32: 60.2%).
  • Contract evidence supports the answer less often than the large model (46% vs 62%); act AUROC on typed-decisions 0.60.
  • The act / escalate thresholds shipped with the model are indicative only: fitted on development data, they are over-confident on new real text (see solvi-large). For a guaranteed risk, calibrate on your own labelled stream with act_guard (solvi ≥ 0.5.1, table above).
  • Several questions per pass disagree with one-question-per-pass in ~10% of answers on real states — disabled by default.
  • English only; truncation beyond 512 tokens (1024 in block layout). Not for medical, legal or credit decisions on its own.

Files

model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]), onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask), onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures, act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format), LICENSES.json (training data manifest).

Citation

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
calibration
classification
decision
distillation
extractive-qa
modernbert
onnx
safetensors
solvi
text-classification
typed-questions