OpenDecision is an open-source semantic decision engine like typesafe's jev.
Python
1
9 commits
updated Sep 19, 2026
OpenDecision is an open-source semantic decision engine.
Give it some state, a natural-language question, and answer criteria. It returns a structured decision.
A few days ago, I saw TypeSafe announce Jev, their first "System One Model". While it looked impressive, it also kind of tingled my spidey sense.
A couple of years ago, I had worked on something similar for a client, using zero-shot models in a narrow domain involving healthcare insurance fraud. I also have experience training and building things with zero-shot models through projects like Zink, so I thought I'd give this a try: build an open-source package that could leverage zero-shot models to provide a similar kind of functionality to what TypeSafe's Jev does.
That experiment became OpenDecision.
OpenDecision currently provides three primitives:
The default backend is MoritzLaurer/ModernBERT-large-zeroshot-v2.0, a compact zero-shot NLI model.
Many application decisions do not require a generative LLM.
Examples:
OpenDecision turns those problems into small, typed semantic decisions.
Python 3.13 is recommended.
git clone https://github.com/deepanwadhwa/OpenDecision.git
cd OpenDecision
uv sync --python 3.13 --all-groups
Start the API:
uv run uvicorn opendecision.api.app:app \
--app-dir src \
--host 127.0.0.1 \
--port 8000
Then open:
http://127.0.0.1:8000/docs
Health check:
curl http://127.0.0.1:8000/health
The model is downloaded from Hugging Face on first use.
from opendecision.engine import OpenDecisionEngine
engine = OpenDecisionEngine()
result = engine.choice(
state="The customer says the router has no power and the ISP line is working.",
instructions="Who should handle this issue?",
criteria={
"isp": "The internet service provider should investigate the connection.",
"router_manufacturer": "The router or its power hardware should be investigated.",
"electric_utility": "The electricity provider should investigate an outage.",
},
)
print(result["choice"])
print(result["probabilities"])
print(result["confidence"])
A Choice response has the form:
{
"type": "choice",
"choice": "router_manufacturer",
"probabilities": {
"isp": ...,
"router_manufacturer": ...,
"electric_utility": ...,
},
"confidence": ...,
}
confidence is a normalized measure of concentration in the returned probability distribution. It is not a calibrated probability that the answer is correct.
result = engine.noul(
state="The request says the production service is currently unavailable.",
instructions="This request is time-sensitive.",
)
print(result["noul"])
result = engine.score(
state="The application crashes for every user at startup.",
instructions="How severe is this software bug?",
criteria={
"0": "Cosmetic or negligible impact.",
"1": "Minor impact with an easy workaround.",
"2": "Significant impact, but the core workflow remains usable.",
"3": "Major failure blocking an important workflow.",
"4": "Critical failure preventing normal use.",
},
)
print(result["score"])
OpenDecision does not rely on one fixed textual formulation for Choice.
The current v0.1 strategy evaluates the request with two complementary semantic compilers:
Compiler A
premise: state
candidates: answer descriptions
hypothesis: question-conditioned
Compiler B
premise: state
candidates: label + description
hypothesis: default NLI hypothesis
If A and B agree, their answer is returned.
If they disagree, OpenDecision runs a small semantic adjudication over the two competing answers:
premise: state + question
candidates: labels only
hypothesis: default NLI hypothesis
This uses the same underlying ModernBERT model throughout; no additional router model is required.
Because OpenDecision was inspired by TypeSafe's Jev, I wanted to see how it performed on tasks resembling the examples TypeSafe itself publishes.
I created an evaluation set of 80 cases adapted from TypeSafe's public documentation and cookbooks:
Using the default MoritzLaurer/ModernBERT-large-zeroshot-v2.0 backend, OpenDecision v0.1 achieves:
| Primitive | Result |
|---|---|
| Choice | 43/51 — 84.3% |
| Noul | 17/20 — 85.0% |
| Score | 0.375 MAE |
These are not official TypeSafe benchmark results and should not be interpreted as a direct Jev-vs-OpenDecision comparison. The cases were adapted and paraphrased from examples in TypeSafe's public documentation so they could be evaluated reproducibly with OpenDecision.
The full evaluation set and source provenance are available in benchmarks/typesafe_public/.
I also created a separate synthetic benchmark containing 500 Choice problems across 25 domains.
The system architecture was developed on 375 cases and then evaluated on a separate 125-case comparison split.
OpenDecision v0.1 scored 108/125 — 86.4%.
The benchmark covers tasks including support and function routing, citation relations, semantic extraction, date semantics, product taxonomy, entity matching, policy decisions, software bugs, scientific methods, security events, and word-sense disambiguation.
The dataset and split are available in benchmarks/opendecision_original/.
OpenDecision exposes a /v1/systemone endpoint compatible with the core request/response flow used by the TypeSafe SDK.
A local OpenDecision server can therefore be used as a base_url for compatible clients.
Once the server is running:
GET /health
POST /v1/systemone
GET /docs
GET /openapi.json
FastAPI's interactive documentation at /docs is the easiest way to inspect the exact request schema and try requests manually.
General benchmark runner:
uv run python benchmarks/run_eval.py \
--cases benchmarks/cases.jsonl
Original benchmark:
uv run python benchmarks/run_eval.py \
--cases benchmarks/opendecision_original/dev.jsonl
The experimental compiler/adjudicator scripts live under benchmarks/.
uv run --python 3.13 pytest -v
CI runs the same test suite on GitHub Actions.
OpenDecision is intentionally small:
Choice path may require three classifier passes when the two primary compilers disagree.v0.1.0 is intended as a developer preview.
The core primitives, local API, SDK-compatible endpoint, evaluation harness, and reproducible benchmark are available for experimentation.
9 commits
Python
100.0%
OpenDecision is an open-source semantic decision engine like typesafe's jev.
Python
1
9 commits
updated Sep 19, 2026
OpenDecision is an open-source semantic decision engine.
Give it some state, a natural-language question, and answer criteria. It returns a structured decision.
A few days ago, I saw TypeSafe announce Jev, their first "System One Model". While it looked impressive, it also kind of tingled my spidey sense.
A couple of years ago, I had worked on something similar for a client, using zero-shot models in a narrow domain involving healthcare insurance fraud. I also have experience training and building things with zero-shot models through projects like Zink, so I thought I'd give this a try: build an open-source package that could leverage zero-shot models to provide a similar kind of functionality to what TypeSafe's Jev does.
That experiment became OpenDecision.
OpenDecision currently provides three primitives:
The default backend is MoritzLaurer/ModernBERT-large-zeroshot-v2.0, a compact zero-shot NLI model.
Many application decisions do not require a generative LLM.
Examples:
OpenDecision turns those problems into small, typed semantic decisions.
Python 3.13 is recommended.
git clone https://github.com/deepanwadhwa/OpenDecision.git
cd OpenDecision
uv sync --python 3.13 --all-groups
Start the API:
uv run uvicorn opendecision.api.app:app \
--app-dir src \
--host 127.0.0.1 \
--port 8000
Then open:
http://127.0.0.1:8000/docs
Health check:
curl http://127.0.0.1:8000/health
The model is downloaded from Hugging Face on first use.
from opendecision.engine import OpenDecisionEngine
engine = OpenDecisionEngine()
result = engine.choice(
state="The customer says the router has no power and the ISP line is working.",
instructions="Who should handle this issue?",
criteria={
"isp": "The internet service provider should investigate the connection.",
"router_manufacturer": "The router or its power hardware should be investigated.",
"electric_utility": "The electricity provider should investigate an outage.",
},
)
print(result["choice"])
print(result["probabilities"])
print(result["confidence"])
A Choice response has the form:
{
"type": "choice",
"choice": "router_manufacturer",
"probabilities": {
"isp": ...,
"router_manufacturer": ...,
"electric_utility": ...,
},
"confidence": ...,
}
confidence is a normalized measure of concentration in the returned probability distribution. It is not a calibrated probability that the answer is correct.
result = engine.noul(
state="The request says the production service is currently unavailable.",
instructions="This request is time-sensitive.",
)
print(result["noul"])
result = engine.score(
state="The application crashes for every user at startup.",
instructions="How severe is this software bug?",
criteria={
"0": "Cosmetic or negligible impact.",
"1": "Minor impact with an easy workaround.",
"2": "Significant impact, but the core workflow remains usable.",
"3": "Major failure blocking an important workflow.",
"4": "Critical failure preventing normal use.",
},
)
print(result["score"])
OpenDecision does not rely on one fixed textual formulation for Choice.
The current v0.1 strategy evaluates the request with two complementary semantic compilers:
Compiler A
premise: state
candidates: answer descriptions
hypothesis: question-conditioned
Compiler B
premise: state
candidates: label + description
hypothesis: default NLI hypothesis
If A and B agree, their answer is returned.
If they disagree, OpenDecision runs a small semantic adjudication over the two competing answers:
premise: state + question
candidates: labels only
hypothesis: default NLI hypothesis
This uses the same underlying ModernBERT model throughout; no additional router model is required.
Because OpenDecision was inspired by TypeSafe's Jev, I wanted to see how it performed on tasks resembling the examples TypeSafe itself publishes.
I created an evaluation set of 80 cases adapted from TypeSafe's public documentation and cookbooks:
Using the default MoritzLaurer/ModernBERT-large-zeroshot-v2.0 backend, OpenDecision v0.1 achieves:
| Primitive | Result |
|---|---|
| Choice | 43/51 — 84.3% |
| Noul | 17/20 — 85.0% |
| Score | 0.375 MAE |
These are not official TypeSafe benchmark results and should not be interpreted as a direct Jev-vs-OpenDecision comparison. The cases were adapted and paraphrased from examples in TypeSafe's public documentation so they could be evaluated reproducibly with OpenDecision.
The full evaluation set and source provenance are available in benchmarks/typesafe_public/.
I also created a separate synthetic benchmark containing 500 Choice problems across 25 domains.
The system architecture was developed on 375 cases and then evaluated on a separate 125-case comparison split.
OpenDecision v0.1 scored 108/125 — 86.4%.
The benchmark covers tasks including support and function routing, citation relations, semantic extraction, date semantics, product taxonomy, entity matching, policy decisions, software bugs, scientific methods, security events, and word-sense disambiguation.
The dataset and split are available in benchmarks/opendecision_original/.
OpenDecision exposes a /v1/systemone endpoint compatible with the core request/response flow used by the TypeSafe SDK.
A local OpenDecision server can therefore be used as a base_url for compatible clients.
Once the server is running:
GET /health
POST /v1/systemone
GET /docs
GET /openapi.json
FastAPI's interactive documentation at /docs is the easiest way to inspect the exact request schema and try requests manually.
General benchmark runner:
uv run python benchmarks/run_eval.py \
--cases benchmarks/cases.jsonl
Original benchmark:
uv run python benchmarks/run_eval.py \
--cases benchmarks/opendecision_original/dev.jsonl
The experimental compiler/adjudicator scripts live under benchmarks/.
uv run --python 3.13 pytest -v
CI runs the same test suite on GitHub Actions.
OpenDecision is intentionally small:
Choice path may require three classifier passes when the two primary compilers disagree.v0.1.0 is intended as a developer preview.
The core primitives, local API, SDK-compatible endpoint, evaluation harness, and reproducible benchmark are available for experimentation.
9 commits
Python
100.0%