LiquidAI/d1-3B

Model

src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"

96

8 commits

updated Oct 7, 2026

See the code

README

Liquid AI
Try LFM β€’ Docs β€’ Discord

d1-3B

d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.

  • Best decision model under 10B on the Decision Index 0.2.1: 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
  • Multimodal: images and text in the same state. It scores 74.1 on 11 public image benchmarks (LFM2.5-VL-3B: 73.9).
  • Fast: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.

Find more information about open d1 in our blog post.

image

[!NOTE] πŸ’» Demos: Try d1-3B in a Hugging Face space without any setup: Open d1 Arcade: Collection of 10 demos using d1-3B

πŸ—’οΈ Model Details

ModelParametersDescription
LFM2.5-VL-3B3.1BGeneral-purpose vision-language model (base)
d1-3B3.1BPost-trained for single-pass, calibrated decisions

d1-3B is a multimodal decision model with the following features:

  • Total parameters: 3.12B
  • Vision encoder: SigLIP2 NaFlex shape-optimized 400M
  • Context length: 32,768 tokens
  • Vocabulary size: 128,000

We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating: routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge scoring, agent guardrails, and visual inspection. It is not a chat model and does not write text.

πŸƒ How to use

Install the dependencies (requires transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

The model ships its own code, so load it with trust_remote_code=True:

import torch
from transformers import AutoModel
from transformers.image_utils import load_image

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
dtype = torch.float32 if device == "cpu" else torch.bfloat16
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=dtype).to(device)

# Text: several named questions over one state, answered in one pass
questions = {
    "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for a refund?",
    },
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "billing": "Charges, refunds, invoices",
            "technical": "App or site faults",
            "fraud": "Suspected unauthorised use",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["Can wait", "Today", "Blocking the customer now"],
    },
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# Image: the photo is the whole state
image = load_image("http://images.cocodataset.org/val2017/000000039769.jpg")  # two cats on a sofa
cats = {
    "type": "choice",
    "instructions": "How many cats are there?",
    "criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}
print(model.system_one(None, {"cats": cats}, images=[image]))

# Batch: many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
call
system_one(state, questions, images=None)Named questions over one state, in one pass. The state and its images are read once for all questions.
system_one_batch([(state, questions[, images]), ...])Many requests, packed with no padding.

A state is a string, any JSON value, or None when the images are the whole state.

Questions and answers

Questions follow the Decision Index schema: type, instructions, and criteria.

typecriteriaanswer fields
noul: yes or nooptional: {"true": "...", "false": "..."} to define each sidenoul: P(yes)
choice: one of named options{name: description}choice, confidence, probabilities
score: 2 to 10 ordered levelsa list of level descriptions, lowest firstscore (the expected level), confidence, probabilities, legend

Each call returns {"answers": {name: answer}, "usage": {"input_tokens": n, "output_tokens": 0}}.

⚑ Speed

Warm calls, one request at a time: a single question, three questions over one state, a 3.4k-token state and a 384 px image. The last column is throughput with 64 states packed into one pass.

Edge Inference

We measure latency on an Apple M5 Pro and, in collaboration with NVIDIA, on an NVIDIA Jetson AGX Thor, a Jetson AGX Orin 64 GB and a Jetson Orin Nano.

one question3 questions, one pass3.4k-token state384 px image64 states, packed
Apple M5 Pro (mps)30 ms41 ms640 ms62 ms78 / s
NVIDIA Jetson AGX Thor16 ms20 ms220 ms35 ms262 / s
NVIDIA Jetson AGX Orin 64 GB26 ms35 ms560 ms83 ms110 / s
NVIDIA Jetson Orin Nano50 ms73 ms1640 ms202 ms38 / s

GPU Inference

We measure latency on an NVIDIA RTX 4090 and an AMD MI325X, in bf16, median of 20 runs.

one question3 questions, one pass3.4k-token state384 px image64 states, packed
NVIDIA RTX 40908 ms21 ms102 ms17 ms475 / s
AMD MI325X9 ms14 ms44 ms18 ms1,106 / s

On NVIDIA GPUs, model.compile(mode="reduce-overhead") runs single questions as CUDA graphs (the RTX 4090 row uses it). Without it, a single question takes 16 ms. The first call with a new shape pays for kernel selection or compilation, so warm up the shapes you serve.

πŸ“Š Performance

All results are on public benchmarks.

Decision Index 0.2.1

We scored d1-3B with the official scorer (not a leaderboard submission). All other rows come from the public leaderboard v0.2.1.

ModelSizeDecision IndexKnowledgeLanguageRetrievalToolsArts
Winnow-12B12B50.0233.856.054.071.030.0
d1-3B3B48.5723.856.452.874.536.3
Decider 35B-A3B36B47.1131.855.554.756.532.6
JPT-9B9.7B46.8931.756.744.667.028.6
Decision 1.0 Lux9.7B43.4930.948.050.057.226.4
JPT-4B4.7B43.0428.752.545.057.225.8
Jet v6.24.7B42.6028.743.948.262.927.0
Decider 4B4.7B40.7025.746.044.758.625.0
Winnow-E4B8.0B39.8922.345.143.862.522.8
Decider 2B2.3B28.9714.932.637.342.414.6

Benchmarks as decisions

Besides the Decision Index, we added a few other internal evaluations based on public benchmarks.

Benchmarkd1-3BDecider 4BDecider 2B
SQuAD 2.085.376.067.7
Civil Comments93.092.893.6
MASSIVE intent87.388.381.1
HelpSteer236.742.032.0
PubMedQA66.063.365.7
BoolQ86.789.087.3
XNLI85.088.685.0
PAWS-X76.969.859.5
Mean77.176.271.5

d1-3B also scores 71.8 on DecisionBench (eng v1, all 23,900 rows) and 69.3 on Fast Decisions (dev split).

Vision

Eleven public image benchmarks, read as decisions over each benchmark's options (at most 1,000 rows each), compared with the base model:

Benchmarkd1-3BLFM2.5-VL-3B
AI2D79.980.9
BLINK59.258.7
CV-Bench82.187.6
HallusionBench65.365.0
MMBench84.984.3
MME82.182.4
MMStar59.961.2
MMVP77.073.7
POPE88.590.1
VisualWebBench71.478.3
VL-RewardBench65.050.9
Mean74.173.9
ImajevBench (dev and calibration, 253 rows)64.066.8

With the images removed, the same questions score 45.1, so the answers come from the images.

πŸ“¬ Contact

Citation

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {https://www.liquid.ai/blog/d1-open},
}
@article{liquidai2025lfm2,
  title   = {LFM2 Technical Report},
  author  = {Liquid AI},
  journal = {arXiv preprint arXiv:2511.23404},
  year    = {2025}
}
calibration
classification
conversational
custom_code
decision
decision-model
edge
endpoints_compatible
image-text-to-text
lfm2.5
lfm2_vl
liquid
multimodal
safetensors
system-one
transformers

LiquidAI/d1-3B

Model

src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"

96

8 commits

updated Oct 7, 2026

See the code

README

Liquid AI
Try LFM β€’ Docs β€’ Discord

d1-3B

d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.

  • Best decision model under 10B on the Decision Index 0.2.1: 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
  • Multimodal: images and text in the same state. It scores 74.1 on 11 public image benchmarks (LFM2.5-VL-3B: 73.9).
  • Fast: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.

Find more information about open d1 in our blog post.

image

[!NOTE] πŸ’» Demos: Try d1-3B in a Hugging Face space without any setup: Open d1 Arcade: Collection of 10 demos using d1-3B

πŸ—’οΈ Model Details

ModelParametersDescription
LFM2.5-VL-3B3.1BGeneral-purpose vision-language model (base)
d1-3B3.1BPost-trained for single-pass, calibrated decisions

d1-3B is a multimodal decision model with the following features:

  • Total parameters: 3.12B
  • Vision encoder: SigLIP2 NaFlex shape-optimized 400M
  • Context length: 32,768 tokens
  • Vocabulary size: 128,000

We recommend d1-3B wherever a pipeline needs a yes/no, a pick from named options, or a rating: routing and triage, moderation, intent and topic classification, extraction checks, reranking, LLM-as-a-judge scoring, agent guardrails, and visual inspection. It is not a chat model and does not write text.

πŸƒ How to use

Install the dependencies (requires transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

The model ships its own code, so load it with trust_remote_code=True:

import torch
from transformers import AutoModel
from transformers.image_utils import load_image

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
dtype = torch.float32 if device == "cpu" else torch.bfloat16
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=dtype).to(device)

# Text: several named questions over one state, answered in one pass
questions = {
    "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for a refund?",
    },
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
            "billing": "Charges, refunds, invoices",
            "technical": "App or site faults",
            "fraud": "Suspected unauthorised use",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["Can wait", "Today", "Blocking the customer now"],
    },
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# Image: the photo is the whole state
image = load_image("http://images.cocodataset.org/val2017/000000039769.jpg")  # two cats on a sofa
cats = {
    "type": "choice",
    "instructions": "How many cats are there?",
    "criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}
print(model.system_one(None, {"cats": cats}, images=[image]))

# Batch: many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
call
system_one(state, questions, images=None)Named questions over one state, in one pass. The state and its images are read once for all questions.
system_one_batch([(state, questions[, images]), ...])Many requests, packed with no padding.

A state is a string, any JSON value, or None when the images are the whole state.

Questions and answers

Questions follow the Decision Index schema: type, instructions, and criteria.

typecriteriaanswer fields
noul: yes or nooptional: {"true": "...", "false": "..."} to define each sidenoul: P(yes)
choice: one of named options{name: description}choice, confidence, probabilities
score: 2 to 10 ordered levelsa list of level descriptions, lowest firstscore (the expected level), confidence, probabilities, legend

Each call returns {"answers": {name: answer}, "usage": {"input_tokens": n, "output_tokens": 0}}.

⚑ Speed

Warm calls, one request at a time: a single question, three questions over one state, a 3.4k-token state and a 384 px image. The last column is throughput with 64 states packed into one pass.

Edge Inference

We measure latency on an Apple M5 Pro and, in collaboration with NVIDIA, on an NVIDIA Jetson AGX Thor, a Jetson AGX Orin 64 GB and a Jetson Orin Nano.

one question3 questions, one pass3.4k-token state384 px image64 states, packed
Apple M5 Pro (mps)30 ms41 ms640 ms62 ms78 / s
NVIDIA Jetson AGX Thor16 ms20 ms220 ms35 ms262 / s
NVIDIA Jetson AGX Orin 64 GB26 ms35 ms560 ms83 ms110 / s
NVIDIA Jetson Orin Nano50 ms73 ms1640 ms202 ms38 / s

GPU Inference

We measure latency on an NVIDIA RTX 4090 and an AMD MI325X, in bf16, median of 20 runs.

one question3 questions, one pass3.4k-token state384 px image64 states, packed
NVIDIA RTX 40908 ms21 ms102 ms17 ms475 / s
AMD MI325X9 ms14 ms44 ms18 ms1,106 / s

On NVIDIA GPUs, model.compile(mode="reduce-overhead") runs single questions as CUDA graphs (the RTX 4090 row uses it). Without it, a single question takes 16 ms. The first call with a new shape pays for kernel selection or compilation, so warm up the shapes you serve.

πŸ“Š Performance

All results are on public benchmarks.

Decision Index 0.2.1

We scored d1-3B with the official scorer (not a leaderboard submission). All other rows come from the public leaderboard v0.2.1.

ModelSizeDecision IndexKnowledgeLanguageRetrievalToolsArts
Winnow-12B12B50.0233.856.054.071.030.0
d1-3B3B48.5723.856.452.874.536.3
Decider 35B-A3B36B47.1131.855.554.756.532.6
JPT-9B9.7B46.8931.756.744.667.028.6
Decision 1.0 Lux9.7B43.4930.948.050.057.226.4
JPT-4B4.7B43.0428.752.545.057.225.8
Jet v6.24.7B42.6028.743.948.262.927.0
Decider 4B4.7B40.7025.746.044.758.625.0
Winnow-E4B8.0B39.8922.345.143.862.522.8
Decider 2B2.3B28.9714.932.637.342.414.6

Benchmarks as decisions

Besides the Decision Index, we added a few other internal evaluations based on public benchmarks.

Benchmarkd1-3BDecider 4BDecider 2B
SQuAD 2.085.376.067.7
Civil Comments93.092.893.6
MASSIVE intent87.388.381.1
HelpSteer236.742.032.0
PubMedQA66.063.365.7
BoolQ86.789.087.3
XNLI85.088.685.0
PAWS-X76.969.859.5
Mean77.176.271.5

d1-3B also scores 71.8 on DecisionBench (eng v1, all 23,900 rows) and 69.3 on Fast Decisions (dev split).

Vision

Eleven public image benchmarks, read as decisions over each benchmark's options (at most 1,000 rows each), compared with the base model:

Benchmarkd1-3BLFM2.5-VL-3B
AI2D79.980.9
BLINK59.258.7
CV-Bench82.187.6
HallusionBench65.365.0
MMBench84.984.3
MME82.182.4
MMStar59.961.2
MMVP77.073.7
POPE88.590.1
VisualWebBench71.478.3
VL-RewardBench65.050.9
Mean74.173.9
ImajevBench (dev and calibration, 253 rows)64.066.8

With the images removed, the same questions score 45.1, so the answers come from the images.

πŸ“¬ Contact

Citation

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {https://www.liquid.ai/blog/d1-open},
}
@article{liquidai2025lfm2,
  title   = {LFM2 Technical Report},
  author  = {Liquid AI},
  journal = {arXiv preprint arXiv:2511.23404},
  year    = {2025}
}
calibration
classification
conversational
custom_code
decision
decision-model
edge
endpoints_compatible
image-text-to-text
lfm2.5
lfm2_vl
liquid
multimodal
safetensors
system-one
transformers