realslapout/arc-1

ARC-1: a fast 1.7B decision model (choice, score, yes/no) with calibrated probabilities

Python

0

3 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

ARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti (r/LocalLLaMA)

I spent the last 11 days training a small model for typed decisions: routing support tickets, intent detection, moderation, "should the agent call this tool", that kind of thing. You give it some context and a question, and it gives you back a choice, a score or a yes/no probability. * 1.7B…

2

Oct 6, 2026

README

ARC-1

ARC-1 is a small decision model. You give it some context and a question, and it picks an option, gives a score or answers yes/no, with a probability for every answer. It has 1.7B parameters, runs on a normal gaming GPU and answers a short request in about 16 ms.

I built it over 11 days on a single RTX 4060 Ti. The question I wanted to answer was how close a small model that runs on your own machine can get to hosted decision APIs like Jev. Short version: it is a lot faster and you can run it offline, but it is not as accurate yet. All the numbers are below, including the ones where it loses.

The ARC-1 demo answering a banking question with eight options

Quick start

pip install git+https://github.com/realslapout/arc-1
from arc1 import ARC1Predictor

model = ARC1Predictor("realslapout/ARC-1")   # downloads ~3.5 GB the first time

state = {"ticket": "I was charged twice for the same order and I want my money back."}
questions = {
    "team": {
        "type": "choice",
        "instructions": "Which team should handle `ticket`?",
        "criteria": {
            "billing": "payments, charges and refunds",
            "shipping": "delivery problems",
            "tech": "app or login problems",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is `ticket`?",
        "criteria": ["not urgent", "normal", "urgent"],
    },
    "angry": {"type": "noul", "instructions": "Is the customer angry?"},
}

answers = model.predict(state, questions)["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgency"]["score"])
print(answers["angry"]["noul"])
billing {'billing': 0.991669, 'shipping': 0.006424, 'tech': 0.001907}
1.0302799518917043
0.9748723399342193

(Measured on a GPU in bf16. On a CPU the numbers come out slightly different, for example billing 0.997.)

Every answer also carries answer_confidence, and score answers include the probability of each level. All questions about the same state are answered in one forward pass, so asking three questions costs not much more than asking one.

For the lowest latency on a GPU, turn on CUDA graphs:

model = ARC1Predictor("realslapout/ARC-1", cuda_graphs=True)

It also runs on a CPU (device="cpu"), at roughly 0.4 s per short request with two threads.

There is also a small web demo (the screenshot above). It runs on your own machine:

pip install gradio
git clone https://github.com/realslapout/arc-1 && cd arc-1
python demo/app.py          # add --share for a temporary public link

Question types

typewhat you give itwhat you get back
choicecriteria: a dict {label: description} or a plain list of labelsthe chosen label and a probability for each label
scorecriteria: a list of levels from lowest to highestthe expected level (a float) and a probability for each level
noulonly instructions (a yes/no statement or question)the probability that the answer is yes

state can be a string, a dict or a list (for example a chat history). Refer to its fields in the instructions with backticks, like `ticket`. The input window is 1,024 tokens. Longer states are cut, and a list keeps its most recent entries. A choice question can have many options: lists that do not fit the window are scored in chunks.

Results

Decision benchmarks that ARC-1 was not trained on

modelsizeJevBench public 231DecideBench v1.1speed
ARC-11.7B, open68.475.525 ms median on an RTX 4060 Ti (local)
Jev 1.13.0undisclosed, API86.698.0~620 ms median (API, network included)
Strands Decider 2B1.9B, open72.3 *–115 ms median on an RTX 3090 *
decider-2b1.9B, open71.0––
Laya421M, open58.459.8–

Accuracy in %. JevBench public 231 is the 231 published JevBench items. The other models' results come from the JevBench repository (results v1.2), and ARC-1 is measured on the same items. DecideBench numbers for other models come from DecideBench, and ARC-1 is measured on our copy of DecideBench v1.1 (400 items). * = reported by the authors.

So, to be clear: Jev is much more accurate, and the two other ~2B models are a few points better on JevBench. What ARC-1 has going for it is speed, running locally, and being free and open.

Held-out items of the training tasks

Most of these datasets, or close relatives of them, are in the training data, so read these as in-domain results on held-out items (anything that overlapped with the evaluation items was removed from training).

taskdatasetmetricARC-1
banking intents (77 options)Banking77accuracy87.2
assistant intents (150 options)CLINC150accuracy92.0
assistant intents, English (60)MASSIVEaccuracy84.2
assistant intents, 16 languages (20 options)MASSIVEaccuracy80.7
news topicAG Newsaccuracy90.6
entity typeDBpedia-14accuracy99.2
natural-language inferenceMNLIaccuracy83.0
adversarial NLIANLIaccuracy52.2
NLI, 12 languagesXNLIaccuracy71.0
yes/no questionsBoolQaccuracy82.6
duplicate questionsQQPbalanced accuracy82.5
paraphrase, 7 languagesPAWS-Xaccuracy83.6
toxic chat messagesToxicChatbalanced accuracy89.4
jailbreak attemptsToxicChatbalanced accuracy94.9
prompt injectiondeepset prompt-injectionsbalanced accuracy82.3
spamEnron spambalanced accuracy91.5
tool selection (2 to 256 tools to choose from)Glaive function calling v2accuracy89.0

One safety set that is not in the training data: on XSTest (safe prompts that look unsafe, plus unsafe contrasts), ARC-1 gets 88.9% balanced accuracy, so it mostly doesn't refuse harmless requests just because of scary words.

Speed

Measured with this package on an RTX 4060 Ti 16 GB, batch size 1, bf16, CUDA graphs on:

requestmedian latency
short request (~60 tokens, 3–4 options)16 ms
JevBench items with 4 options25 ms (90th percentile 100 ms)
full 1,024-token input~100 ms

Latency grows with the input length because the model reads the whole state on every request. GPU memory use is about 4 GB.

How it works

  • Backbone: Qwen3-1.7B-Base, fine-tuned with LoRA (rank 64) and merged into the weights.
  • Layout: the state and the question are encoded once. Every option then gets its own short branch that can see the state and the question but not the other options, so the order of the options does not matter. A small head scores each branch. This started from the input format of Laya, but replaces its encoder with a decoder backbone and adds a token budget, so long option lists are not cut down to a couple of tokens each.
  • Calibration: one temperature per question type, fitted on dev data, so the probabilities mean roughly what they say.
  • Training: about 3.8 million examples from around 330 public datasets, covering intents, topics, sentiment, inference and logic, moderation and safety, prompt injection, support tickets, tool selection and preference data, plus generated decision tasks. The full list with licence tags is in DATA.md. Training rows that overlapped any evaluation item were removed before training.
  • The released checkpoint is "milestone 8": an interpolation of two training runs (run 6 plus 0.3 of the difference to run 9). The interpolation kept most of run 9's gains while undoing an over-cautious safety behaviour that run 9 had picked up.

Limitations

  • It is much less accurate than large hosted models on hard, multi-step decisions (the JevBench "hard" items).
  • The input window is 1,024 tokens.
  • It was trained mostly on English. Other languages work, but less well (see the MASSIVE and XNLI rows).
  • The probabilities are calibrated on our dev data. If you rely on the thresholds, check them on your own data.
  • Don't use it for high-stakes decisions (medical, legal, financial, hiring and so on) without a human checking the result.

License

  • Code (this repository): Apache 2.0, see LICENSE and NOTICE.
  • Model weights: CC BY-NC 4.0, so non-commercial use only. About 7% of the training examples come from datasets that only allow non-commercial use (PKU-SafeRLHF, ANLI, BeaverTails, ToxicChat and others, listed in DATA.md), and many others don't state a licence, so the weights can't be offered for commercial use.

Thanks

To the Qwen team for the base model, to Convai and Nandha Kishor M for Laya, to the authors of the datasets in DATA.md (a lot of them reached me through the tasksource collection), and to the people who maintain JevBench and DecideBench.

classification
decision-making
guardrails
intent-detection
llm
qwen3
small-language-model
zero-shot-classification

realslapout/arc-1

ARC-1: a fast 1.7B decision model (choice, score, yes/no) with calibrated probabilities

Python

0

3 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

ARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti (r/LocalLLaMA)

I spent the last 11 days training a small model for typed decisions: routing support tickets, intent detection, moderation, "should the agent call this tool", that kind of thing. You give it some context and a question, and it gives you back a choice, a score or a yes/no probability. * 1.7B…

2

Oct 6, 2026

README

ARC-1

ARC-1 is a small decision model. You give it some context and a question, and it picks an option, gives a score or answers yes/no, with a probability for every answer. It has 1.7B parameters, runs on a normal gaming GPU and answers a short request in about 16 ms.

I built it over 11 days on a single RTX 4060 Ti. The question I wanted to answer was how close a small model that runs on your own machine can get to hosted decision APIs like Jev. Short version: it is a lot faster and you can run it offline, but it is not as accurate yet. All the numbers are below, including the ones where it loses.

The ARC-1 demo answering a banking question with eight options

Quick start

pip install git+https://github.com/realslapout/arc-1
from arc1 import ARC1Predictor

model = ARC1Predictor("realslapout/ARC-1")   # downloads ~3.5 GB the first time

state = {"ticket": "I was charged twice for the same order and I want my money back."}
questions = {
    "team": {
        "type": "choice",
        "instructions": "Which team should handle `ticket`?",
        "criteria": {
            "billing": "payments, charges and refunds",
            "shipping": "delivery problems",
            "tech": "app or login problems",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is `ticket`?",
        "criteria": ["not urgent", "normal", "urgent"],
    },
    "angry": {"type": "noul", "instructions": "Is the customer angry?"},
}

answers = model.predict(state, questions)["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgency"]["score"])
print(answers["angry"]["noul"])
billing {'billing': 0.991669, 'shipping': 0.006424, 'tech': 0.001907}
1.0302799518917043
0.9748723399342193

(Measured on a GPU in bf16. On a CPU the numbers come out slightly different, for example billing 0.997.)

Every answer also carries answer_confidence, and score answers include the probability of each level. All questions about the same state are answered in one forward pass, so asking three questions costs not much more than asking one.

For the lowest latency on a GPU, turn on CUDA graphs:

model = ARC1Predictor("realslapout/ARC-1", cuda_graphs=True)

It also runs on a CPU (device="cpu"), at roughly 0.4 s per short request with two threads.

There is also a small web demo (the screenshot above). It runs on your own machine:

pip install gradio
git clone https://github.com/realslapout/arc-1 && cd arc-1
python demo/app.py          # add --share for a temporary public link

Question types

typewhat you give itwhat you get back
choicecriteria: a dict {label: description} or a plain list of labelsthe chosen label and a probability for each label
scorecriteria: a list of levels from lowest to highestthe expected level (a float) and a probability for each level
noulonly instructions (a yes/no statement or question)the probability that the answer is yes

state can be a string, a dict or a list (for example a chat history). Refer to its fields in the instructions with backticks, like `ticket`. The input window is 1,024 tokens. Longer states are cut, and a list keeps its most recent entries. A choice question can have many options: lists that do not fit the window are scored in chunks.

Results

Decision benchmarks that ARC-1 was not trained on

modelsizeJevBench public 231DecideBench v1.1speed
ARC-11.7B, open68.475.525 ms median on an RTX 4060 Ti (local)
Jev 1.13.0undisclosed, API86.698.0~620 ms median (API, network included)
Strands Decider 2B1.9B, open72.3 *–115 ms median on an RTX 3090 *
decider-2b1.9B, open71.0––
Laya421M, open58.459.8–

Accuracy in %. JevBench public 231 is the 231 published JevBench items. The other models' results come from the JevBench repository (results v1.2), and ARC-1 is measured on the same items. DecideBench numbers for other models come from DecideBench, and ARC-1 is measured on our copy of DecideBench v1.1 (400 items). * = reported by the authors.

So, to be clear: Jev is much more accurate, and the two other ~2B models are a few points better on JevBench. What ARC-1 has going for it is speed, running locally, and being free and open.

Held-out items of the training tasks

Most of these datasets, or close relatives of them, are in the training data, so read these as in-domain results on held-out items (anything that overlapped with the evaluation items was removed from training).

taskdatasetmetricARC-1
banking intents (77 options)Banking77accuracy87.2
assistant intents (150 options)CLINC150accuracy92.0
assistant intents, English (60)MASSIVEaccuracy84.2
assistant intents, 16 languages (20 options)MASSIVEaccuracy80.7
news topicAG Newsaccuracy90.6
entity typeDBpedia-14accuracy99.2
natural-language inferenceMNLIaccuracy83.0
adversarial NLIANLIaccuracy52.2
NLI, 12 languagesXNLIaccuracy71.0
yes/no questionsBoolQaccuracy82.6
duplicate questionsQQPbalanced accuracy82.5
paraphrase, 7 languagesPAWS-Xaccuracy83.6
toxic chat messagesToxicChatbalanced accuracy89.4
jailbreak attemptsToxicChatbalanced accuracy94.9
prompt injectiondeepset prompt-injectionsbalanced accuracy82.3
spamEnron spambalanced accuracy91.5
tool selection (2 to 256 tools to choose from)Glaive function calling v2accuracy89.0

One safety set that is not in the training data: on XSTest (safe prompts that look unsafe, plus unsafe contrasts), ARC-1 gets 88.9% balanced accuracy, so it mostly doesn't refuse harmless requests just because of scary words.

Speed

Measured with this package on an RTX 4060 Ti 16 GB, batch size 1, bf16, CUDA graphs on:

requestmedian latency
short request (~60 tokens, 3–4 options)16 ms
JevBench items with 4 options25 ms (90th percentile 100 ms)
full 1,024-token input~100 ms

Latency grows with the input length because the model reads the whole state on every request. GPU memory use is about 4 GB.

How it works

  • Backbone: Qwen3-1.7B-Base, fine-tuned with LoRA (rank 64) and merged into the weights.
  • Layout: the state and the question are encoded once. Every option then gets its own short branch that can see the state and the question but not the other options, so the order of the options does not matter. A small head scores each branch. This started from the input format of Laya, but replaces its encoder with a decoder backbone and adds a token budget, so long option lists are not cut down to a couple of tokens each.
  • Calibration: one temperature per question type, fitted on dev data, so the probabilities mean roughly what they say.
  • Training: about 3.8 million examples from around 330 public datasets, covering intents, topics, sentiment, inference and logic, moderation and safety, prompt injection, support tickets, tool selection and preference data, plus generated decision tasks. The full list with licence tags is in DATA.md. Training rows that overlapped any evaluation item were removed before training.
  • The released checkpoint is "milestone 8": an interpolation of two training runs (run 6 plus 0.3 of the difference to run 9). The interpolation kept most of run 9's gains while undoing an over-cautious safety behaviour that run 9 had picked up.

Limitations

  • It is much less accurate than large hosted models on hard, multi-step decisions (the JevBench "hard" items).
  • The input window is 1,024 tokens.
  • It was trained mostly on English. Other languages work, but less well (see the MASSIVE and XNLI rows).
  • The probabilities are calibrated on our dev data. If you rely on the thresholds, check them on your own data.
  • Don't use it for high-stakes decisions (medical, legal, financial, hiring and so on) without a human checking the result.

License

  • Code (this repository): Apache 2.0, see LICENSE and NOTICE.
  • Model weights: CC BY-NC 4.0, so non-commercial use only. About 7% of the training examples come from datasets that only allow non-commercial use (PKU-SafeRLHF, ANLI, BeaverTails, ToxicChat and others, listed in DATA.md), and many others don't state a licence, so the weights can't be offered for commercial use.

Thanks

To the Qwen team for the base model, to Convai and Nandha Kishor M for Laya, to the authors of the datasets in DATA.md (a lot of them reached me through the tasksource collection), and to the people who maintain JevBench and DecideBench.

classification
decision-making
guardrails
intent-detection
llm
qwen3
small-language-model
zero-shot-classification