Cloudflare/clef

Model

Clef

71

10 commits

updated Oct 1, 2026

See the code

README

Clef

Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing.

The Clef API is fully compatible with Jev and SystemOne.

Clef is post-trained from Qwen/Qwen3.8-27B. See Clef-Flash for the smaller, faster variant.

Model

  • Backbone: Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
  • Joint schema head: a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.
  • Output: one logit per allowed option for each question. Apply a softmax per question to get probabilities.

Files

FilePurpose
model-*.safetensors, model.safetensors.index.json, config.json, generation_config.jsonBackbone, including the vision encoder
joint_head.safetensors, joint_head_config.jsonJoint schema head
joint_schema_model.pyRecord encoding, batching, the model, load_release_model, and systemone
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.jsonTokenizer and image/video processor
LICENSEApache-2.0 license

Usage

Tested with torch 2.11 and transformers 5.10.2 on a single H200. Image and video inputs also need pillow.

import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

Jev / SystemOne API

systemone takes a Jev/SystemOne POST /v1/systemone request body and returns the same response body: model, answers keyed by question ID, and usage. A choice answer has choice, confidence, and probabilities; a score answer has the expected score, confidence, legend, and probabilities; a noul answer has the probability of true. instructions is optional, and images and videos may be added to the request.

from joint_schema_model import systemone

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

Images and video

Add images (PIL images) or videos (frame arrays) to the record and pass the processor to encode_record. Optional processor arguments go in media_kwargs.

from PIL import Image

record = {
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)

Text-only and multimodal records can be mixed in the same batch.

Input format

FieldDescription
stateAny string or JSON value describing the situation to decide on
images, videosOptional lists of images or video frame arrays
media_kwargsOptional keyword arguments for the image/video processor
questionsMapping of question ID to question

Each question has:

  • type: noul (true/false), choice (named options), or score (ordered options)
  • instructions: what to decide; optional, and the question ID is used when it is omitted
  • criteria: for choice, a mapping of option ID to description; for score, a list of option descriptions indexed from 0; for noul, optional descriptions for true and false

encode_record accepts max_length (default 16,384 tokens) and max_state_tokens to bound the input.

Results

Decision Index

Per-benchmark results from our internal run of the Decision Index 0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.

BenchmarkClefClef-flashJevDiffusionGemma JevKev 9BLaya
BFCL (case exact accuracy)98.598.895.896.594.538.1
ToolRet (nDCG@10)69.266.465.361.264.312.8
API-Bank (accuracy)91.993.188.283.756.311.5
BANKING77 (macro-F1)94.290.979.774.384.814.3
CLINC150+OOS (macro-F1)97.466.889.383.579.03.2
RouterBench (selected quality)79.779.979.979.080.057.1
Home appliance simulator (case exact accuracy)83.097.752.342.025.00.0
SGD/SGD-X (macro-F1)43.834.243.040.664.042.4
ContractNLI (macro-F1)81.484.371.776.057.829.0
ANLI (macro-F1)69.859.174.866.456.348.7
BPoMP (accuracy)96.995.490.686.967.051.6
Humicroedit (accuracy)66.775.161.963.055.847.2
POP909-CL (accuracy)15.81.618.12.510.85.1
cfcolor (accuracy)66.065.864.758.256.352.3
MMLU (accuracy)90.391.891.779.375.330.7
GPQA Diamond (accuracy)48.051.078.344.938.827.6
ARC-Easy (accuracy)99.099.599.398.297.747.0
ARC-Challenge (accuracy)97.798.397.894.593.728.6
WinoGrande (accuracy)93.597.592.073.673.250.5
HellaSwag (accuracy)98.298.694.583.381.933.1
GSM8K (accuracy)80.867.379.950.348.721.6
ChessBench (accuracy)24.723.017.214.211.27.7
MuSR (accuracy)83.586.066.161.257.943.2
SATA-Bench (case exact accuracy)33.836.726.427.526.70.3
BRIGHT (nDCG@10)45.939.347.542.938.519.9
Amazon ESCI (macro-F1)57.557.455.253.449.224.4
ACOS (per-review F1)33.325.929.524.518.33.5
FinEntity (macro-F1)96.297.187.089.088.461.0
VAST (macro-F1)59.549.664.655.755.440.5
NLI4CT (macro-F1)82.978.684.178.474.947.7
CRUXEval (accuracy)86.786.173.064.751.240.2
CLadder (accuracy)94.097.772.667.862.052.9
ForecastBench (Brier, lower is better)13.910.617.429.617.641.1
Habermas Machine (accuracy)68.771.845.945.039.433.4
PhishNChips (accuracy)79.675.062.585.450.750.1
MMLU-Pro (accuracy)65.965.382.756.951.113.6
BBH (accuracy)73.768.992.970.765.234.1
RAGTruth (hallucination F1)79.435.676.570.446.248.8
HoVer (accuracy)65.261.272.970.958.855.8
When2Call MCQ (accuracy)72.465.681.075.449.611.9
New Yorker (accuracy)69.566.170.163.658.127.1
Median latency (ms)209.338.8524.184.451.45.8
p95 latency (ms)238.6122.4536.0211.2187.9222.5

Workflow evals

Decision accuracy on four end-to-end business workflows from Typesafe Evals, scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.

WorkflowMetricClefClef-flashJev
Invoice processingExact actions64.757.161.8
Invoice processingPrimary action86.273.383.1
Customer serviceExact actions76.377.076.0
Security incidentsExact actions62.961.761.7
Agent trace observabilityPrimary action68.569.871.6

License

Released under the Apache-2.0 license, following the base model Qwen/Qwen3.8-27B.

classification
clef
cloudflare
conversational
custom-code
endpoints_compatible
image-text-to-text
image-text-to-typed-output
multimodal
post-train
qwen3_5
qwen3.8
safetensors
structured-output
systemone
transformers

Cloudflare/clef

Model

Clef

71

10 commits

updated Oct 1, 2026

See the code

README

Clef

Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing.

The Clef API is fully compatible with Jev and SystemOne.

Clef is post-trained from Qwen/Qwen3.8-27B. See Clef-Flash for the smaller, faster variant.

Model

  • Backbone: Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
  • Joint schema head: a small transformer head that reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.
  • Output: one logit per allowed option for each question. Apply a softmax per question to get probabilities.

Files

FilePurpose
model-*.safetensors, model.safetensors.index.json, config.json, generation_config.jsonBackbone, including the vision encoder
joint_head.safetensors, joint_head_config.jsonJoint schema head
joint_schema_model.pyRecord encoding, batching, the model, load_release_model, and systemone
tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.jsonTokenizer and image/video processor
LICENSEApache-2.0 license

Usage

Tested with torch 2.11 and transformers 5.10.2 on a single H200. Image and video inputs also need pillow.

import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

Jev / SystemOne API

systemone takes a Jev/SystemOne POST /v1/systemone request body and returns the same response body: model, answers keyed by question ID, and usage. A choice answer has choice, confidence, and probabilities; a score answer has the expected score, confidence, legend, and probabilities; a noul answer has the probability of true. instructions is optional, and images and videos may be added to the request.

from joint_schema_model import systemone

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

Images and video

Add images (PIL images) or videos (frame arrays) to the record and pass the processor to encode_record. Optional processor arguments go in media_kwargs.

from PIL import Image

record = {
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)

Text-only and multimodal records can be mixed in the same batch.

Input format

FieldDescription
stateAny string or JSON value describing the situation to decide on
images, videosOptional lists of images or video frame arrays
media_kwargsOptional keyword arguments for the image/video processor
questionsMapping of question ID to question

Each question has:

  • type: noul (true/false), choice (named options), or score (ordered options)
  • instructions: what to decide; optional, and the question ID is used when it is omitted
  • criteria: for choice, a mapping of option ID to description; for score, a list of option descriptions indexed from 0; for noul, optional descriptions for true and false

encode_record accepts max_length (default 16,384 tokens) and max_state_tokens to bound the input.

Results

Decision Index

Per-benchmark results from our internal run of the Decision Index 0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.

BenchmarkClefClef-flashJevDiffusionGemma JevKev 9BLaya
BFCL (case exact accuracy)98.598.895.896.594.538.1
ToolRet (nDCG@10)69.266.465.361.264.312.8
API-Bank (accuracy)91.993.188.283.756.311.5
BANKING77 (macro-F1)94.290.979.774.384.814.3
CLINC150+OOS (macro-F1)97.466.889.383.579.03.2
RouterBench (selected quality)79.779.979.979.080.057.1
Home appliance simulator (case exact accuracy)83.097.752.342.025.00.0
SGD/SGD-X (macro-F1)43.834.243.040.664.042.4
ContractNLI (macro-F1)81.484.371.776.057.829.0
ANLI (macro-F1)69.859.174.866.456.348.7
BPoMP (accuracy)96.995.490.686.967.051.6
Humicroedit (accuracy)66.775.161.963.055.847.2
POP909-CL (accuracy)15.81.618.12.510.85.1
cfcolor (accuracy)66.065.864.758.256.352.3
MMLU (accuracy)90.391.891.779.375.330.7
GPQA Diamond (accuracy)48.051.078.344.938.827.6
ARC-Easy (accuracy)99.099.599.398.297.747.0
ARC-Challenge (accuracy)97.798.397.894.593.728.6
WinoGrande (accuracy)93.597.592.073.673.250.5
HellaSwag (accuracy)98.298.694.583.381.933.1
GSM8K (accuracy)80.867.379.950.348.721.6
ChessBench (accuracy)24.723.017.214.211.27.7
MuSR (accuracy)83.586.066.161.257.943.2
SATA-Bench (case exact accuracy)33.836.726.427.526.70.3
BRIGHT (nDCG@10)45.939.347.542.938.519.9
Amazon ESCI (macro-F1)57.557.455.253.449.224.4
ACOS (per-review F1)33.325.929.524.518.33.5
FinEntity (macro-F1)96.297.187.089.088.461.0
VAST (macro-F1)59.549.664.655.755.440.5
NLI4CT (macro-F1)82.978.684.178.474.947.7
CRUXEval (accuracy)86.786.173.064.751.240.2
CLadder (accuracy)94.097.772.667.862.052.9
ForecastBench (Brier, lower is better)13.910.617.429.617.641.1
Habermas Machine (accuracy)68.771.845.945.039.433.4
PhishNChips (accuracy)79.675.062.585.450.750.1
MMLU-Pro (accuracy)65.965.382.756.951.113.6
BBH (accuracy)73.768.992.970.765.234.1
RAGTruth (hallucination F1)79.435.676.570.446.248.8
HoVer (accuracy)65.261.272.970.958.855.8
When2Call MCQ (accuracy)72.465.681.075.449.611.9
New Yorker (accuracy)69.566.170.163.658.127.1
Median latency (ms)209.338.8524.184.451.45.8
p95 latency (ms)238.6122.4536.0211.2187.9222.5

Workflow evals

Decision accuracy on four end-to-end business workflows from Typesafe Evals, scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.

WorkflowMetricClefClef-flashJev
Invoice processingExact actions64.757.161.8
Invoice processingPrimary action86.273.383.1
Customer serviceExact actions76.377.076.0
Security incidentsExact actions62.961.761.7
Agent trace observabilityPrimary action68.569.871.6

License

Released under the Apache-2.0 license, following the base model Qwen/Qwen3.8-27B.

classification
clef
cloudflare
conversational
custom-code
endpoints_compatible
image-text-to-text
image-text-to-typed-output
multimodal
post-train
qwen3_5
qwen3.8
safetensors
structured-output
systemone
transformers