illate/qwen3-4b-banking77-lora

Model

Qwen3-4B Banking77 LoRA

0

4 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen3-4B Banking77 LoRA

A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that sorts bank customer messages into the 77 Banking77 intents. Built by ILLATE as a public example of a parity check: does a small open model you own match the job, measured on the full test set?

Results (full test set, 3,080 messages)

ModelAccuracy95% CIMacro-F1Invalid outputs
TF-IDF + logistic regression (free baseline)89.3%88.1% to 90.3%89.3%0
This adapter94.0%93.1% to 94.8%94.0%0
Qwen3-4B, same model, no fine-tuning (served, label list in prompt)63.1%
  • Difference against the baseline: +4.7 points (paired bootstrap 95% CI +3.7 to +5.8).
  • Latency on one NVIDIA L4: p50 43 ms, p95 60 ms per message (one at a time).
  • Served with vLLM on one L4: 93.9% accuracy, about 87 requests/s at p95 0.55 s.
  • Confidence intervals are bootstrap resamples over the test messages.

Most common mistakes are near-duplicates in the label set, for example fiat_currency_support read as exchange_via_app (5 times) and card_arrival as card_delivery_estimate (4 times).

Training

  • Data: Banking77 training set (CC BY 4.0, PolyAI), 9,000 messages for training and 1,003 held out as a dev set; the official 3,080-message test set is untouched.
  • LoRA rank 32, alpha 32, dropout 0.05, on all attention and MLP projections.
  • 2 epochs, learning rate 2e-4 with cosine schedule, batch 16, bf16, one NVIDIA L4 (about 22 minutes).
  • Labels are human-written dataset labels; no outputs from closed models were used.

How to use

The model answers with the label name. Use the same system prompt as in training:

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "illate/qwen3-4b-banking77-lora")

messages = [
    {"role": "system", "content": "Classify the bank customer's message. Reply with the category name only."},
    {"role": "user", "content": "I still haven't received my new card, when will it arrive?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=16, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))  # card_arrival

The reported numbers use constrained decoding, which only lets the model write one of the 77 label names. Without it, the model can occasionally produce a label that does not exist; with vLLM, pass the label list as a structured-output choice.

With vLLM:

vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --lora-modules banking77=illate/qwen3-4b-banking77-lora

Limits

  • Trained for Banking77's 77 intents only; it is not a general banking assistant.
  • English only.
  • Benchmark data is public and short (one message each). Accuracy on your own traffic will differ; measure it before relying on it.

License

Adapter: Apache 2.0, same as the base model. Training data: Banking77, CC BY 4.0 (Casanueva et al., 2020, "Efficient Intent Detection with Dual Sentence Encoders").

Contact: poojith@illate.dev · illate.dev

banking77
conversational
illate
intent-classification
lora
model-index
peft
safetensors
text-classification
text-generation

illate/qwen3-4b-banking77-lora

Model

Qwen3-4B Banking77 LoRA

0

4 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen3-4B Banking77 LoRA

A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that sorts bank customer messages into the 77 Banking77 intents. Built by ILLATE as a public example of a parity check: does a small open model you own match the job, measured on the full test set?

Results (full test set, 3,080 messages)

ModelAccuracy95% CIMacro-F1Invalid outputs
TF-IDF + logistic regression (free baseline)89.3%88.1% to 90.3%89.3%0
This adapter94.0%93.1% to 94.8%94.0%0
Qwen3-4B, same model, no fine-tuning (served, label list in prompt)63.1%
  • Difference against the baseline: +4.7 points (paired bootstrap 95% CI +3.7 to +5.8).
  • Latency on one NVIDIA L4: p50 43 ms, p95 60 ms per message (one at a time).
  • Served with vLLM on one L4: 93.9% accuracy, about 87 requests/s at p95 0.55 s.
  • Confidence intervals are bootstrap resamples over the test messages.

Most common mistakes are near-duplicates in the label set, for example fiat_currency_support read as exchange_via_app (5 times) and card_arrival as card_delivery_estimate (4 times).

Training

  • Data: Banking77 training set (CC BY 4.0, PolyAI), 9,000 messages for training and 1,003 held out as a dev set; the official 3,080-message test set is untouched.
  • LoRA rank 32, alpha 32, dropout 0.05, on all attention and MLP projections.
  • 2 epochs, learning rate 2e-4 with cosine schedule, batch 16, bf16, one NVIDIA L4 (about 22 minutes).
  • Labels are human-written dataset labels; no outputs from closed models were used.

How to use

The model answers with the label name. Use the same system prompt as in training:

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "illate/qwen3-4b-banking77-lora")

messages = [
    {"role": "system", "content": "Classify the bank customer's message. Reply with the category name only."},
    {"role": "user", "content": "I still haven't received my new card, when will it arrive?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=16, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))  # card_arrival

The reported numbers use constrained decoding, which only lets the model write one of the 77 label names. Without it, the model can occasionally produce a label that does not exist; with vLLM, pass the label list as a structured-output choice.

With vLLM:

vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --lora-modules banking77=illate/qwen3-4b-banking77-lora

Limits

  • Trained for Banking77's 77 intents only; it is not a general banking assistant.
  • English only.
  • Benchmark data is public and short (one message each). Accuracy on your own traffic will differ; measure it before relying on it.

License

Adapter: Apache 2.0, same as the base model. Training data: Banking77, CC BY 4.0 (Casanueva et al., 2020, "Efficient Intent Detection with Dual Sentence Encoders").

Contact: poojith@illate.dev · illate.dev

banking77
conversational
illate
intent-classification
lora
model-index
peft
safetensors
text-classification
text-generation