Evaluate and harden type-safe generation around TypeLLM
See the codeKind: landing page (the project overview; the one page that mixes kinds).
typevet is a Python library that asks a model typed questions and returns typed answers.
The three question types are Noul (yes or no), Choice (one label) and Score (one rubric level).
typevet computes each answer from the model's next-token probabilities, read before sampling.
typevet also returns JSON objects that pass a JSON Schema you supply, or it raises an error.
The receipts cover Gemma 4 31B on llama.cpp for local work and on vLLM for hosting.
Read the documentation at https://alberto-codes.github.io/typevet/.
0.1.0.| Backend | Tested pin | Receipt |
|---|---|---|
| vLLM | vllm/vllm-openai:v0.30.0, BF16 google/gemma-4-31B-it, one H100 80 GB | #170 |
| llama.cpp | Build b11223-4da633776, local alias gemma-4-31b-kv9-q4km-mm | #203 |
| llama.cpp grammar | Build b11243-fc07d781e, Gemma 4 31B QAT Q4_0 GGUF | #129 |
Performance: on one H100 at concurrency level 64, 480 Banking77 records took 12.1 s at 39.6 records/s. That run had 0 errors. Banking77 calibration passed; DIFrauD SMS failed parity (ECE 0.158 against 0.10). One run, one pod, one pin. See Performance on one H100 and Serve Gemma 4 31B on a rented H100. A valid structure does not prove accuracy or calibration. The receipts are small samples.
Get one offline typed judgment from a scripted fake. This step needs no model.
uv sync
uv run python -c "
from typevet.domain import Noul
from typevet.runtime import ScoringJudgmentAdapter
from typevet.testing import ScriptedScoringFake
fake = ScriptedScoringFake(logprobs={'True': -0.2, 'False': -1.0})
port = ScoringJudgmentAdapter(fake, tokenize_content=lambda t: (ord(t[0]),))
r = port.judge('text', {'q': Noul(instructions='Ok?', criteria={'true': 'Y', 'false': 'N'})}, 'fake')
print('noul', r.nouls['q'].noul)
"
The command prints the probability of yes, near 0.69. The offline tutorial explains each step. Then connect a model server:
TypeLLM is a research reference for the decision model. It is not a runtime dependency.
Read CLAUDE.md first. It states the gates, the issue workflow and the rules for agents and people.
uv sync
uv run pre-commit install -t pre-commit -t pre-push -t commit-msg
uv run pytest -q
The default test run skips live tests.
Pull requests and pushes to main run the hook stages in
the CI workflow.
The writing system and
the commit rules apply to every change.
Python
99.6%
Evaluate and harden type-safe generation around TypeLLM
See the codeKind: landing page (the project overview; the one page that mixes kinds).
typevet is a Python library that asks a model typed questions and returns typed answers.
The three question types are Noul (yes or no), Choice (one label) and Score (one rubric level).
typevet computes each answer from the model's next-token probabilities, read before sampling.
typevet also returns JSON objects that pass a JSON Schema you supply, or it raises an error.
The receipts cover Gemma 4 31B on llama.cpp for local work and on vLLM for hosting.
Read the documentation at https://alberto-codes.github.io/typevet/.
0.1.0.| Backend | Tested pin | Receipt |
|---|---|---|
| vLLM | vllm/vllm-openai:v0.30.0, BF16 google/gemma-4-31B-it, one H100 80 GB | #170 |
| llama.cpp | Build b11223-4da633776, local alias gemma-4-31b-kv9-q4km-mm | #203 |
| llama.cpp grammar | Build b11243-fc07d781e, Gemma 4 31B QAT Q4_0 GGUF | #129 |
Performance: on one H100 at concurrency level 64, 480 Banking77 records took 12.1 s at 39.6 records/s. That run had 0 errors. Banking77 calibration passed; DIFrauD SMS failed parity (ECE 0.158 against 0.10). One run, one pod, one pin. See Performance on one H100 and Serve Gemma 4 31B on a rented H100. A valid structure does not prove accuracy or calibration. The receipts are small samples.
Get one offline typed judgment from a scripted fake. This step needs no model.
uv sync
uv run python -c "
from typevet.domain import Noul
from typevet.runtime import ScoringJudgmentAdapter
from typevet.testing import ScriptedScoringFake
fake = ScriptedScoringFake(logprobs={'True': -0.2, 'False': -1.0})
port = ScoringJudgmentAdapter(fake, tokenize_content=lambda t: (ord(t[0]),))
r = port.judge('text', {'q': Noul(instructions='Ok?', criteria={'true': 'Y', 'false': 'N'})}, 'fake')
print('noul', r.nouls['q'].noul)
"
The command prints the probability of yes, near 0.69. The offline tutorial explains each step. Then connect a model server:
TypeLLM is a research reference for the decision model. It is not a runtime dependency.
Read CLAUDE.md first. It states the gates, the issue workflow and the rules for agents and people.
uv sync
uv run pre-commit install -t pre-commit -t pre-push -t commit-msg
uv run pytest -q
The default test run skips live tests.
Pull requests and pushes to main run the hook stages in
the CI workflow.
The writing system and
the commit rules apply to every change.
Python
99.6%