A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that sorts bank customer messages into the 77 Banking77 intents. Built by ILLATE as a public example of a parity check: does a small open model you own match the job, measured on the full test set?
| Model | Accuracy | 95% CI | Macro-F1 | Invalid outputs |
|---|---|---|---|---|
| TF-IDF + logistic regression (free baseline) | 89.3% | 88.1% to 90.3% | 89.3% | 0 |
| This adapter | 94.0% | 93.1% to 94.8% | 94.0% | 0 |
| Qwen3-4B, same model, no fine-tuning (served, label list in prompt) | 63.1% |
Most common mistakes are near-duplicates in the label set, for example
fiat_currency_support read as exchange_via_app (5 times) and card_arrival as
card_delivery_estimate (4 times).
The model answers with the label name. Use the same system prompt as in training:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "illate/qwen3-4b-banking77-lora")
messages = [
{"role": "system", "content": "Classify the bank customer's message. Reply with the category name only."},
{"role": "user", "content": "I still haven't received my new card, when will it arrive?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=16, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)) # card_arrival
The reported numbers use constrained decoding, which only lets the model write one of
the 77 label names. Without it, the model can occasionally produce a label that does
not exist; with vLLM, pass the label list as a structured-output choice.
With vLLM:
vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --lora-modules banking77=illate/qwen3-4b-banking77-lora
Adapter: Apache 2.0, same as the base model. Training data: Banking77, CC BY 4.0 (Casanueva et al., 2020, "Efficient Intent Detection with Dual Sentence Encoders").
Contact: poojith@illate.dev · illate.dev
A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that sorts bank customer messages into the 77 Banking77 intents. Built by ILLATE as a public example of a parity check: does a small open model you own match the job, measured on the full test set?
| Model | Accuracy | 95% CI | Macro-F1 | Invalid outputs |
|---|---|---|---|---|
| TF-IDF + logistic regression (free baseline) | 89.3% | 88.1% to 90.3% | 89.3% | 0 |
| This adapter | 94.0% | 93.1% to 94.8% | 94.0% | 0 |
| Qwen3-4B, same model, no fine-tuning (served, label list in prompt) | 63.1% |
Most common mistakes are near-duplicates in the label set, for example
fiat_currency_support read as exchange_via_app (5 times) and card_arrival as
card_delivery_estimate (4 times).
The model answers with the label name. Use the same system prompt as in training:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "illate/qwen3-4b-banking77-lora")
messages = [
{"role": "system", "content": "Classify the bank customer's message. Reply with the category name only."},
{"role": "user", "content": "I still haven't received my new card, when will it arrive?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=16, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)) # card_arrival
The reported numbers use constrained decoding, which only lets the model write one of
the 77 label names. Without it, the model can occasionally produce a label that does
not exist; with vLLM, pass the label list as a structured-output choice.
With vLLM:
vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --lora-modules banking77=illate/qwen3-4b-banking77-lora
Adapter: Apache 2.0, same as the base model. Training data: Banking77, CC BY 4.0 (Casanueva et al., 2020, "Efficient Intent Detection with Dual Sentence Encoders").
Contact: poojith@illate.dev · illate.dev