
Website • Demo • GitHub • PyPI
DecisionTune 1.0 is a 395M decision model. You give it a state, a question and a list of options. It selects one option and gives a probability for each option. You can also ask a yes/no question. Then it gives P(yes). It does one encoder pass and does not generate text. It scores 29.57 on Decision Index 0.2.1. We measured all numbers on our runs. Estimates have a label.
Key features

Weights: Apache-2.0. Package: decision-tune (source). Demo: decision-tune/demo. Site: decisiontune.com.
[!TIP] Describe your options. Short descriptions ("Shipping: delivery, lost or damaged packages") give much better results than labels only ("shipping").
The package is small. Before the first download of the weights (1.58 GB), it asks for your approval. It keeps the files in your Hugging Face cache. Then it checks each file against manifest.json.
pip and Python
pip install decision-tune
from decision_tune import DecisionModel
m = DecisionModel.from_pretrained("decision-tune/decisiontune-1.0")
m.choose("The order arrived broken.", "Which team should handle this customer message?",
{"billing": "Billing: charges, refunds, invoices",
"shipping": "Shipping: delivery, lost or damaged packages",
"tech": "Tech support: bugs, crashes, login problems"})
# {'choice': 'shipping', 'probabilities': {'billing': ..., 'shipping': ..., 'tech': ...}, 'confidence': ...}
m.yes_no("The order arrived broken. I want my money back.", "Is the customer asking for a refund?")
# P(yes): here, a float near 1
Command line (with uvx, you do not need to install)
uvx decision-tune ask "Is the customer asking for a refund?" --state "The order arrived broken. I want my money back."
decision-tune ask "Which team handles this?" --state "I was charged twice this month." --option billing --option shipping --option "tech support"
decision-tune download --yes # download without a prompt, for scripts and containers
Local HTTP endpoint
decision-tune serve # http://127.0.0.1:8000/decide
In a second terminal:
curl -s http://127.0.0.1:8000/decide -H 'Content-Type: application/json' -d '{"state": "What is the weather tomorrow in Paris?", "question": "Which tool should be called?", "options": ["get_weather", "send_email", "create_calendar_event"]}'
curl -s http://127.0.0.1:8000/decide -H 'Content-Type: application/json' -d '{"state": "The order arrived broken. I want my money back.", "question": "Is the customer asking for a refund?"}'
Local app. Both command names work: decision-tune and decisiontune.
curl -LsSf https://decisiontune.com/install.sh | sh # or: uv tool install "decision-tune[mlx]" on a Mac
decisiontune app # opens the app in your browser
The app runs on your computer. Nothing leaves it. On a Mac with Apple silicon (MLX), a short decision takes about 10 ms. If port 8000 is busy, the app tries the next 10 ports.
Recipes (many items at once). A recipe is a JSON file with the columns to read and the questions to ask. support-triage is built in. Each result row has needs_review. It is true for low confidence or an error. Output files put an apostrophe before any cell that starts with =, +, - or @.
decisiontune run support-triage tickets.csv # writes tickets-decided.csv; use -o out.xlsx for Excel
decisiontune recipes # list the recipes
decisiontune recipe new my-recipe # save a copy to edit
from decision_tune import Recipe
results = Recipe.load("support-triage").run([{"subject": "Broken mug", "message": "The order arrived broken."}])
MCP (Claude Desktop, Cursor and others). Run decisiontune download once. Then add this to your client settings:
{"mcpServers": {"decisiontune": {"command": "decisiontune", "args": ["mcp"]}}}
The server has three tools: decide, run_recipe and list_recipes. It runs on your computer. Add --allow FOLDER to args to limit the assistant to that folder; without it the tools can read any file you can.
Backends. PyTorch is the default. On Apple silicon, pip install "decision-tune[mlx]" adds an MLX backend. The package selects it automatically. The MLX backend uses laya-mlx (Apache-2.0). pip install "decision-tune[onnx]" adds onnxruntime. To use it, set --backend onnx or backend="onnx". All three backends read the files in this repo. They give the same answers on our parity rows.
Python without the package. engine.py in this repo is all of the inference code.
pip install torch transformers tokenizers safetensors huggingface_hub
hf download decision-tune/decisiontune-1.0 --exclude "onnx/*" --local-dir decisiontune-1.0
cd decisiontune-1.0
python -c 'from engine import DecisionModel; print(DecisionModel(".").yes_no("The order arrived broken. I want my money back.", "Is the customer asking for a refund?"))'

The model is ModernBERT-large and a 4 KB scoring head. Each question is one sequence:
[CLS] question [SEP] [MASK] option_1 [MASK] option_2 ... [SEP] state [SEP]
The head gives a score to the hidden state at each [MASK]. The probabilities are the softmax of these scores. A yes/no question uses the options yes and no. There is no calibration temperature and no option filter. The context limit is 8,192 tokens. The model refuses a longer sequence. It does not truncate it. A state can be a string or a JSON object. The model shows an object as key: value | key: value.
Decision Index 0.2.1, one full run: 29.57 (raw 46.80). The run had 150,317 requests that the index can score. 150,309 were ok and 8 were not supported. The context limit is 8,192 tokens.
| Area (skill) | DecisionTune 1.0 | DecisionTune 0.9 Preview | Change |
|---|---|---|---|
| Knowledge & Reasoning | 13.3 | 12.2 | +1.1 |
| Language Understanding | 31.5 | 29.0 | +2.5 |
| Retrieval & Classification | 45.0 | 44.8 | +0.2 |
| Tools & Automation | 46.5 | 28.1 | +18.4 |
| Arts & Human Taste | 4.8 | 3.7 | +1.1 |
Bold: the best number in each row.
DecisionTune 0.9 Preview is our earlier clean model. It scores 25.12. Nine benchmarks score 0.0. Examples are ANLI, GPQA Diamond, ChessBench and HLE.
Speed. In the index run, the median time for one request was 25.9 ms, and p95 was 407.8 ms (our laptop). Speed on your hardware can be different.
Parity. Torch and MLX give the same answer on 99.85% of 2,755 questions. Before each release, the package runs 50 recorded index rows again on each backend. PyTorch, ONNX and MLX each give the recorded answer on all 50.
Use it to route, sort and classify short English states with a small number of options. Use it to select tools. Use it for yes/no checks when you want a probability. It is small, and it runs on a laptop CPU.
The weights have three stages:
We made the tool-use sets, with the When2Call-style rows, from Glaive, ToolACE, Hermes and Funcdex tool dialogs. We did not use NVIDIA When2Call or xLAM data in training. A larger open model, Clef-Flash (Cloudflare/clef-flash, Apache-2.0), taught the model during part of its training. We used its soft labels on 8 datasets (distillation weight 0.5, temperature 1.0).
We fine-tuned DecisionTune 1.0 from answerdotai/ModernBERT-large (Apache-2.0). We trained it only on data that has a permissive license tag from its publisher (Apache-2.0, MIT, CC BY, CC0). We did not use datasets with non-commercial, share-alike or research-only terms. We did not audit the rights of the original text inside the tagged datasets.
Sources: BANKING77 (CC BY 4.0), CLINC150 (CC BY 3.0), GSM8K (MIT), WinoGrande (CC BY), ContractNLI (CC BY 4.0), Amazon ESCI (Apache-2.0), twitter-financial-news-sentiment (MIT), HellaSwag (MIT), RAGTruth train split (MIT), 2WikiMultihopQA (Apache-2.0), glaive-function-calling-v2 (Apache-2.0), ToolACE (Apache-2.0), hermes-function-calling-v1 (Apache-2.0), Funcdex-MT (MIT), procedural-typed-decisions (Apache-2.0), Circa (CC BY 4.0), CondaQA (Apache-2.0), GoEmotions (Apache-2.0), civil_comments (CC0), PubMedQA (MIT), Pavlick formality scores (CC BY 3.0), QMSum (MIT). We made the tool-use, HoVer-style and CLadder-style training sets from these sources and from the CLadder generator code (MIT). We did not use published CLadder rows. Part of the training used soft labels from Clef-Flash (Cloudflare/clef-flash, Apache-2.0).
Known issues. (1) Some sources contain text from other origins. We did not examine the terms of these origins: Wikipedia passages (CondaQA, 2WikiMultihopQA), Reddit comments (GoEmotions), tweets (twitter-financial-news-sentiment), PubMed abstracts (PubMedQA), parliamentary transcripts (QMSum), and web-found sentences and contracts (Pavlick formality, ContractNLI). (2) LLMs made some sources. Some of these LLMs have no name or belong to third parties. RAGTruth responses come from GPT-3.5, GPT-4, Llama-2 and Mistral models. LLM pipelines made the tool-call sets. We used the license tag of each publisher. We did not find the terms of the LLMs. (3) The MIT license link of HellaSwag goes to a source repository that is not available now. (4) CC BY sources require attribution: BANKING77 (PolyAI), CLINC150 (Larson et al.), WinoGrande (Allen Institute for AI), ContractNLI (Koreeda and Manning), Circa (Google Research), Pavlick formality scores (Pavlick and Tetreault). This text describes our data policy. It is not legal advice. Nobody did a combined legal review of the release.
manifest.json lists the SHA-256 of each model file. The decision-tune package contains the SHA-256 of manifest.json. It does not load a file that does not match. The GitHub Actions release workflow also signs the manifest with Sigstore (keyless). The bundle is attached to the v1.0.0 release:
pip install sigstore
python -m sigstore verify github manifest.json --bundle manifest.json.sigstore.json \
--cert-identity https://github.com/decision-tune/decision-tune/.github/workflows/release.yml@refs/tags/v1.0.0
@misc{decisiontune2026,
title = {DecisionTune 1.0: a 395M decision model},
author = {{DecisionTune}},
year = {2026},
howpublished = {\url{https://huggingface.co/decision-tune/decisiontune-1.0}},
note = {hotin.ai}
}

Website • Demo • GitHub • PyPI
DecisionTune 1.0 is a 395M decision model. You give it a state, a question and a list of options. It selects one option and gives a probability for each option. You can also ask a yes/no question. Then it gives P(yes). It does one encoder pass and does not generate text. It scores 29.57 on Decision Index 0.2.1. We measured all numbers on our runs. Estimates have a label.
Key features

Weights: Apache-2.0. Package: decision-tune (source). Demo: decision-tune/demo. Site: decisiontune.com.
[!TIP] Describe your options. Short descriptions ("Shipping: delivery, lost or damaged packages") give much better results than labels only ("shipping").
The package is small. Before the first download of the weights (1.58 GB), it asks for your approval. It keeps the files in your Hugging Face cache. Then it checks each file against manifest.json.
pip and Python
pip install decision-tune
from decision_tune import DecisionModel
m = DecisionModel.from_pretrained("decision-tune/decisiontune-1.0")
m.choose("The order arrived broken.", "Which team should handle this customer message?",
{"billing": "Billing: charges, refunds, invoices",
"shipping": "Shipping: delivery, lost or damaged packages",
"tech": "Tech support: bugs, crashes, login problems"})
# {'choice': 'shipping', 'probabilities': {'billing': ..., 'shipping': ..., 'tech': ...}, 'confidence': ...}
m.yes_no("The order arrived broken. I want my money back.", "Is the customer asking for a refund?")
# P(yes): here, a float near 1
Command line (with uvx, you do not need to install)
uvx decision-tune ask "Is the customer asking for a refund?" --state "The order arrived broken. I want my money back."
decision-tune ask "Which team handles this?" --state "I was charged twice this month." --option billing --option shipping --option "tech support"
decision-tune download --yes # download without a prompt, for scripts and containers
Local HTTP endpoint
decision-tune serve # http://127.0.0.1:8000/decide
In a second terminal:
curl -s http://127.0.0.1:8000/decide -H 'Content-Type: application/json' -d '{"state": "What is the weather tomorrow in Paris?", "question": "Which tool should be called?", "options": ["get_weather", "send_email", "create_calendar_event"]}'
curl -s http://127.0.0.1:8000/decide -H 'Content-Type: application/json' -d '{"state": "The order arrived broken. I want my money back.", "question": "Is the customer asking for a refund?"}'
Local app. Both command names work: decision-tune and decisiontune.
curl -LsSf https://decisiontune.com/install.sh | sh # or: uv tool install "decision-tune[mlx]" on a Mac
decisiontune app # opens the app in your browser
The app runs on your computer. Nothing leaves it. On a Mac with Apple silicon (MLX), a short decision takes about 10 ms. If port 8000 is busy, the app tries the next 10 ports.
Recipes (many items at once). A recipe is a JSON file with the columns to read and the questions to ask. support-triage is built in. Each result row has needs_review. It is true for low confidence or an error. Output files put an apostrophe before any cell that starts with =, +, - or @.
decisiontune run support-triage tickets.csv # writes tickets-decided.csv; use -o out.xlsx for Excel
decisiontune recipes # list the recipes
decisiontune recipe new my-recipe # save a copy to edit
from decision_tune import Recipe
results = Recipe.load("support-triage").run([{"subject": "Broken mug", "message": "The order arrived broken."}])
MCP (Claude Desktop, Cursor and others). Run decisiontune download once. Then add this to your client settings:
{"mcpServers": {"decisiontune": {"command": "decisiontune", "args": ["mcp"]}}}
The server has three tools: decide, run_recipe and list_recipes. It runs on your computer. Add --allow FOLDER to args to limit the assistant to that folder; without it the tools can read any file you can.
Backends. PyTorch is the default. On Apple silicon, pip install "decision-tune[mlx]" adds an MLX backend. The package selects it automatically. The MLX backend uses laya-mlx (Apache-2.0). pip install "decision-tune[onnx]" adds onnxruntime. To use it, set --backend onnx or backend="onnx". All three backends read the files in this repo. They give the same answers on our parity rows.
Python without the package. engine.py in this repo is all of the inference code.
pip install torch transformers tokenizers safetensors huggingface_hub
hf download decision-tune/decisiontune-1.0 --exclude "onnx/*" --local-dir decisiontune-1.0
cd decisiontune-1.0
python -c 'from engine import DecisionModel; print(DecisionModel(".").yes_no("The order arrived broken. I want my money back.", "Is the customer asking for a refund?"))'

The model is ModernBERT-large and a 4 KB scoring head. Each question is one sequence:
[CLS] question [SEP] [MASK] option_1 [MASK] option_2 ... [SEP] state [SEP]
The head gives a score to the hidden state at each [MASK]. The probabilities are the softmax of these scores. A yes/no question uses the options yes and no. There is no calibration temperature and no option filter. The context limit is 8,192 tokens. The model refuses a longer sequence. It does not truncate it. A state can be a string or a JSON object. The model shows an object as key: value | key: value.
Decision Index 0.2.1, one full run: 29.57 (raw 46.80). The run had 150,317 requests that the index can score. 150,309 were ok and 8 were not supported. The context limit is 8,192 tokens.
| Area (skill) | DecisionTune 1.0 | DecisionTune 0.9 Preview | Change |
|---|---|---|---|
| Knowledge & Reasoning | 13.3 | 12.2 | +1.1 |
| Language Understanding | 31.5 | 29.0 | +2.5 |
| Retrieval & Classification | 45.0 | 44.8 | +0.2 |
| Tools & Automation | 46.5 | 28.1 | +18.4 |
| Arts & Human Taste | 4.8 | 3.7 | +1.1 |
Bold: the best number in each row.
DecisionTune 0.9 Preview is our earlier clean model. It scores 25.12. Nine benchmarks score 0.0. Examples are ANLI, GPQA Diamond, ChessBench and HLE.
Speed. In the index run, the median time for one request was 25.9 ms, and p95 was 407.8 ms (our laptop). Speed on your hardware can be different.
Parity. Torch and MLX give the same answer on 99.85% of 2,755 questions. Before each release, the package runs 50 recorded index rows again on each backend. PyTorch, ONNX and MLX each give the recorded answer on all 50.
Use it to route, sort and classify short English states with a small number of options. Use it to select tools. Use it for yes/no checks when you want a probability. It is small, and it runs on a laptop CPU.
The weights have three stages:
We made the tool-use sets, with the When2Call-style rows, from Glaive, ToolACE, Hermes and Funcdex tool dialogs. We did not use NVIDIA When2Call or xLAM data in training. A larger open model, Clef-Flash (Cloudflare/clef-flash, Apache-2.0), taught the model during part of its training. We used its soft labels on 8 datasets (distillation weight 0.5, temperature 1.0).
We fine-tuned DecisionTune 1.0 from answerdotai/ModernBERT-large (Apache-2.0). We trained it only on data that has a permissive license tag from its publisher (Apache-2.0, MIT, CC BY, CC0). We did not use datasets with non-commercial, share-alike or research-only terms. We did not audit the rights of the original text inside the tagged datasets.
Sources: BANKING77 (CC BY 4.0), CLINC150 (CC BY 3.0), GSM8K (MIT), WinoGrande (CC BY), ContractNLI (CC BY 4.0), Amazon ESCI (Apache-2.0), twitter-financial-news-sentiment (MIT), HellaSwag (MIT), RAGTruth train split (MIT), 2WikiMultihopQA (Apache-2.0), glaive-function-calling-v2 (Apache-2.0), ToolACE (Apache-2.0), hermes-function-calling-v1 (Apache-2.0), Funcdex-MT (MIT), procedural-typed-decisions (Apache-2.0), Circa (CC BY 4.0), CondaQA (Apache-2.0), GoEmotions (Apache-2.0), civil_comments (CC0), PubMedQA (MIT), Pavlick formality scores (CC BY 3.0), QMSum (MIT). We made the tool-use, HoVer-style and CLadder-style training sets from these sources and from the CLadder generator code (MIT). We did not use published CLadder rows. Part of the training used soft labels from Clef-Flash (Cloudflare/clef-flash, Apache-2.0).
Known issues. (1) Some sources contain text from other origins. We did not examine the terms of these origins: Wikipedia passages (CondaQA, 2WikiMultihopQA), Reddit comments (GoEmotions), tweets (twitter-financial-news-sentiment), PubMed abstracts (PubMedQA), parliamentary transcripts (QMSum), and web-found sentences and contracts (Pavlick formality, ContractNLI). (2) LLMs made some sources. Some of these LLMs have no name or belong to third parties. RAGTruth responses come from GPT-3.5, GPT-4, Llama-2 and Mistral models. LLM pipelines made the tool-call sets. We used the license tag of each publisher. We did not find the terms of the LLMs. (3) The MIT license link of HellaSwag goes to a source repository that is not available now. (4) CC BY sources require attribution: BANKING77 (PolyAI), CLINC150 (Larson et al.), WinoGrande (Allen Institute for AI), ContractNLI (Koreeda and Manning), Circa (Google Research), Pavlick formality scores (Pavlick and Tetreault). This text describes our data policy. It is not legal advice. Nobody did a combined legal review of the release.
manifest.json lists the SHA-256 of each model file. The decision-tune package contains the SHA-256 of manifest.json. It does not load a file that does not match. The GitHub Actions release workflow also signs the manifest with Sigstore (keyless). The bundle is attached to the v1.0.0 release:
pip install sigstore
python -m sigstore verify github manifest.json --bundle manifest.json.sigstore.json \
--cert-identity https://github.com/decision-tune/decision-tune/.github/workflows/release.yml@refs/tags/v1.0.0
@misc{decisiontune2026,
title = {DecisionTune 1.0: a 395M decision model},
author = {{DecisionTune}},
year = {2026},
howpublished = {\url{https://huggingface.co/decision-tune/decisiontune-1.0}},
note = {hotin.ai}
}