Alberto-Codes/typevet

Evaluate and harden type-safe generation around TypeLLM

Python

0

188 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included (r/LocalLLaMA)

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card,…

0

Sep 29, 2026

README

CI Docs Python Ruff docs vetted

typevet

Kind: landing page (the project overview; the one page that mixes kinds).

typevet is a Python library that asks a model typed questions and returns typed answers. The three question types are Noul (yes or no), Choice (one label) and Score (one rubric level). typevet computes each answer from the model's next-token probabilities, read before sampling. typevet also returns JSON objects that pass a JSON Schema you supply, or it raises an error. The receipts cover Gemma 4 31B on llama.cpp for local work and on vLLM for hosting.

Read the documentation at https://alberto-codes.github.io/typevet/.

Status

  • typevet is pre-1.0. The package version is 0.1.0.
  • typevet is not on PyPI yet. Build a wheel from a checkout and install it: see Install typevet.
  • typevet requires Python 3.12 or later.
  • Each backend has one tested model pin. The receipts give the full pin and its limits.
BackendTested pinReceipt
vLLMvllm/vllm-openai:v0.30.0, BF16 google/gemma-4-31B-it, one H100 80 GB#170
llama.cppBuild b11223-4da633776, local alias gemma-4-31b-kv9-q4km-mm#203
llama.cpp grammarBuild b11243-fc07d781e, Gemma 4 31B QAT Q4_0 GGUF#129

Performance: on one H100 at concurrency level 64, 480 Banking77 records took 12.1 s at 39.6 records/s. That run had 0 errors. Banking77 calibration passed; DIFrauD SMS failed parity (ECE 0.158 against 0.10). One run, one pod, one pin. See Performance on one H100 and Serve Gemma 4 31B on a rented H100. A valid structure does not prove accuracy or calibration. The receipts are small samples.

Quickstart

Get one offline typed judgment from a scripted fake. This step needs no model.

uv sync
uv run python -c "
from typevet.domain import Noul
from typevet.runtime import ScoringJudgmentAdapter
from typevet.testing import ScriptedScoringFake
fake = ScriptedScoringFake(logprobs={'True': -0.2, 'False': -1.0})
port = ScoringJudgmentAdapter(fake, tokenize_content=lambda t: (ord(t[0]),))
r = port.judge('text', {'q': Noul(instructions='Ok?', criteria={'true': 'Y', 'false': 'N'})}, 'fake')
print('noul', r.nouls['q'].noul)
"

The command prints the probability of yes, near 0.69. The offline tutorial explains each step. Then connect a model server:

Learn more

TypeLLM is a research reference for the decision model. It is not a runtime dependency.

For contributors

Read CLAUDE.md first. It states the gates, the issue workflow and the rules for agents and people.

uv sync
uv run pre-commit install -t pre-commit -t pre-push -t commit-msg
uv run pytest -q

The default test run skips live tests. Pull requests and pushes to main run the hook stages in the CI workflow. The writing system and the commit rules apply to every change.

Alberto-Codes/typevet

Evaluate and harden type-safe generation around TypeLLM

Python

0

188 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included (r/LocalLLaMA)

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card,…

0

Sep 29, 2026

README

CI Docs Python Ruff docs vetted

typevet

Kind: landing page (the project overview; the one page that mixes kinds).

typevet is a Python library that asks a model typed questions and returns typed answers. The three question types are Noul (yes or no), Choice (one label) and Score (one rubric level). typevet computes each answer from the model's next-token probabilities, read before sampling. typevet also returns JSON objects that pass a JSON Schema you supply, or it raises an error. The receipts cover Gemma 4 31B on llama.cpp for local work and on vLLM for hosting.

Read the documentation at https://alberto-codes.github.io/typevet/.

Status

  • typevet is pre-1.0. The package version is 0.1.0.
  • typevet is not on PyPI yet. Build a wheel from a checkout and install it: see Install typevet.
  • typevet requires Python 3.12 or later.
  • Each backend has one tested model pin. The receipts give the full pin and its limits.
BackendTested pinReceipt
vLLMvllm/vllm-openai:v0.30.0, BF16 google/gemma-4-31B-it, one H100 80 GB#170
llama.cppBuild b11223-4da633776, local alias gemma-4-31b-kv9-q4km-mm#203
llama.cpp grammarBuild b11243-fc07d781e, Gemma 4 31B QAT Q4_0 GGUF#129

Performance: on one H100 at concurrency level 64, 480 Banking77 records took 12.1 s at 39.6 records/s. That run had 0 errors. Banking77 calibration passed; DIFrauD SMS failed parity (ECE 0.158 against 0.10). One run, one pod, one pin. See Performance on one H100 and Serve Gemma 4 31B on a rented H100. A valid structure does not prove accuracy or calibration. The receipts are small samples.

Quickstart

Get one offline typed judgment from a scripted fake. This step needs no model.

uv sync
uv run python -c "
from typevet.domain import Noul
from typevet.runtime import ScoringJudgmentAdapter
from typevet.testing import ScriptedScoringFake
fake = ScriptedScoringFake(logprobs={'True': -0.2, 'False': -1.0})
port = ScoringJudgmentAdapter(fake, tokenize_content=lambda t: (ord(t[0]),))
r = port.judge('text', {'q': Noul(instructions='Ok?', criteria={'true': 'Y', 'false': 'N'})}, 'fake')
print('noul', r.nouls['q'].noul)
"

The command prints the probability of yes, near 0.69. The offline tutorial explains each step. Then connect a model server:

Learn more

TypeLLM is a research reference for the decision model. It is not a runtime dependency.

For contributors

Read CLAUDE.md first. It states the gates, the issue workflow and the rules for agents and people.

uv sync
uv run pre-commit install -t pre-commit -t pre-push -t commit-msg
uv run pytest -q

The default test run skips live tests. Pull requests and pushes to main run the hook stages in the CI workflow. The writing system and the commit rules apply to every change.

Languages

Python

99.6%