A reasoning-efficiency fine-tune of deepreinforce-ai/Ornith-1.0-9B.
This model is the result of an effort to improve a small coding model by reducing the token count for reasoning traces while retaining accuracy. It worked surprisingly well - 74% shorter reasoning tokens on math prompts. Smaller reductions to actual code, but accuracy also improved.
This is a reproduction of the ThinkingCap method (BottleCap AI) by doing GRPO training with correctness reinforcement and length penalized.
I ran this tune using Claude Code driving the training based on a question - can we apply ThinkingCap to coding models and improve their performance?
This is a LoRA adapter (222 MiB), not a full model — apply it on top of the base
deepreinforce-ai/Ornith-1.0-9B. Want a ready-to-run quantized build instead? See the GGUF
companion repo khudgins/Ornith-1.0-9B-ThinkingCap-GGUF (Q8_0 recommended).
Quick use — transformers + PEFT:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE = "deepreinforce-ai/Ornith-1.0-9B"
ADAPTER = "khudgins/Ornith-1.0-9B-ThinkingCap" # this repo
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER).eval()
msgs = [{"role": "user", "content": "A store had 48 apples, sold 3/4, then got 30 more. How many now?"}]
inp = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",
return_dict=True).to(model.device)
out = model.generate(**inp, max_new_tokens=1024, do_sample=False)
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))
Hosting — vLLM (LoRA served on the base, hot-swappable):
vllm serve deepreinforce-ai/Ornith-1.0-9B \
--enable-lora \
--lora-modules thinking-cap=khudgins/Ornith-1.0-9B-ThinkingCap \
--max-lora-rank 32
# then call the OpenAI-compatible API with "model": "thinking-cap"
Local / quantized — GGUF (Ollama or llama.cpp):
Prebuilt GGUFs (f16 + Q8_0) are in the companion repo
khudgins/Ornith-1.0-9B-ThinkingCap-GGUF.
# Ollama — pulls the GGUF straight from HF
ollama run hf.co/khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 "your prompt"
# llama.cpp — run, or serve an OpenAI-compatible endpoint
llama-cli -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 -p "your prompt"
llama-server -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0
For a permanent single-model HF/vLLM deployment you can instead merge the adapter into the
base (PeftModel.merge_and_unload()) — but serving the adapter keeps the base swappable and the
download tiny. Building your own GGUF from a merged model? Set mtp_num_hidden_layers=0 in its
config.json first, or the converter counts a phantom 33rd block and the GGUF won't load.
Greedy decode, evaluated in-process (transformers + PEFT). Each cell is
accuracy @ mean completion tokens.
| Benchmark (N) | Base Ornith-9B | Thinking-Cap (this model) |
|---|---|---|
| GSM8K (150) | 88.0% @ 533 | 91.3% @ 136 (−74% tokens) |
| HumanEval (164) | 92.7% @ 728 | 92.1% @ 669 (−8% tokens) |
| MBPP (100) | 71.0% @ 846 | 73.0% @ 689 (−19% tokens) |
| average accuracy | 83.9% | 85.5% |
| total tokens | 2107 | 1494 (−29%) |
Adaptive concision. The token cut scales to how compressible each task is — dramatic on math (reasoning is cheap), modest on code.
A frequent failure mode of length-penalized RL is repetitive degeneration ("anxiety") in agentic use. Tested on 24 loop-prone prompts (hard reasoning, eval-framing, ambiguous), greedy decode, with a repeated-span detector:
| Model | Loop rate | Mean tokens | Runaway completions |
|---|---|---|---|
| Base Ornith-9B | 0.0% | 1014 | 3 |
| Thinking-Cap | 0.0% | 607 | 1 |
Neither loops; the tune produces fewer runaway completions than the base model. The correctness reinforcement kept the accuracy up.
GGUF quants preserve the behavior. Benchmarked via Ollama (one harness for both, so the difference isolates the quantization effect; absolute numbers run slightly different from the transformers table above because of the GGUF chat-template path):
| Format | Size | GSM8K | HumanEval | MBPP | avg |
|---|---|---|---|---|---|
| f16 | 17 GB | 91.3% | 90.9% | 77.0% | 86.4% |
| Q8_0 | 9.5 GB | 90.0% | 89.0% | 76.0% | 85.0% |
Q8 lands within 1–3 problems of f16 on every benchmark (−1.4 pts average) with identical token behavior — quantization is effectively free. Q8_0 at half the memory is the recommended daily driver; f16 is the publish-precision reference. (Lower quants are not recommended: a Q4 of the base model failed to complete the eval at all.)
Correctness-gated, length-penalized GRPO (RL with verifiable rewards). Per completion:
reward = correctness # 1.0 if the answer/tests pass, else 0.0
− (λ · normalized_length if correct) # length penalty ONLY when correct
+ format_bonus # well-formed <think>…</think> + answer
The gate is the safety property. The length penalty applies only to correct answers, so the model can never trade accuracy for brevity. GRPO's group-relative advantage means compression only happens on prompts the model has already mastered; a KL leash to the frozen base model and a λ ramp (off → max over ~300 steps) keep it stable and prevent degeneration. The result is self-limiting: the length curve hits a floor and plateaus (checkpoints at steps 400–800 are a tied plateau).
Trained with LoRA (r32/a64, all-linear, MoE router frozen) via Unsloth + TRL GRPOTrainer
on an NVIDIA DGX Spark (GB10). Reward signal reuses the same verifiable checkers used to grade
the model, so training and evaluation share one oracle.
Full recipe, reward harness, and eval scripts: github.com/khudgins/ornith-thinking-cap. Method:
ThinkingCap (correctness-gated length-penalized GRPO) — an independent reproduction of the
ThinkingCap model series by BottleCap AI
(announcement ·
bottlecapai/ThinkingCap-Qwen3.6-27B),
applied here to a smaller coding model. Not affiliated with or endorsed by BottleCap AI. Base model:
deepreinforce-ai/Ornith-1.0-9B (MIT per its HF tag), itself post-trained on Qwen3.5 (Apache
2.0). See License for the full chain. Tooling: Unsloth, TRL, llama.cpp, Ollama.
MIT. This fine-tune — the LoRA adapter and any weights merged from it — is released under the MIT License. It adds no new base weights; it's a low-rank delta trained on top of the base.
Upstream provenance and terms, stated precisely:
deepreinforce-ai/Ornith-1.0-9B
is tagged MIT on Hugging Face; the inherited base weights are governed by that grant.Qwen3_5ForConditionalGeneration).5 commits
A reasoning-efficiency fine-tune of deepreinforce-ai/Ornith-1.0-9B.
This model is the result of an effort to improve a small coding model by reducing the token count for reasoning traces while retaining accuracy. It worked surprisingly well - 74% shorter reasoning tokens on math prompts. Smaller reductions to actual code, but accuracy also improved.
This is a reproduction of the ThinkingCap method (BottleCap AI) by doing GRPO training with correctness reinforcement and length penalized.
I ran this tune using Claude Code driving the training based on a question - can we apply ThinkingCap to coding models and improve their performance?
This is a LoRA adapter (222 MiB), not a full model — apply it on top of the base
deepreinforce-ai/Ornith-1.0-9B. Want a ready-to-run quantized build instead? See the GGUF
companion repo khudgins/Ornith-1.0-9B-ThinkingCap-GGUF (Q8_0 recommended).
Quick use — transformers + PEFT:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE = "deepreinforce-ai/Ornith-1.0-9B"
ADAPTER = "khudgins/Ornith-1.0-9B-ThinkingCap" # this repo
tok = AutoTokenizer.from_pretrained(BASE)
model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER).eval()
msgs = [{"role": "user", "content": "A store had 48 apples, sold 3/4, then got 30 more. How many now?"}]
inp = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",
return_dict=True).to(model.device)
out = model.generate(**inp, max_new_tokens=1024, do_sample=False)
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))
Hosting — vLLM (LoRA served on the base, hot-swappable):
vllm serve deepreinforce-ai/Ornith-1.0-9B \
--enable-lora \
--lora-modules thinking-cap=khudgins/Ornith-1.0-9B-ThinkingCap \
--max-lora-rank 32
# then call the OpenAI-compatible API with "model": "thinking-cap"
Local / quantized — GGUF (Ollama or llama.cpp):
Prebuilt GGUFs (f16 + Q8_0) are in the companion repo
khudgins/Ornith-1.0-9B-ThinkingCap-GGUF.
# Ollama — pulls the GGUF straight from HF
ollama run hf.co/khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 "your prompt"
# llama.cpp — run, or serve an OpenAI-compatible endpoint
llama-cli -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 -p "your prompt"
llama-server -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0
For a permanent single-model HF/vLLM deployment you can instead merge the adapter into the
base (PeftModel.merge_and_unload()) — but serving the adapter keeps the base swappable and the
download tiny. Building your own GGUF from a merged model? Set mtp_num_hidden_layers=0 in its
config.json first, or the converter counts a phantom 33rd block and the GGUF won't load.
Greedy decode, evaluated in-process (transformers + PEFT). Each cell is
accuracy @ mean completion tokens.
| Benchmark (N) | Base Ornith-9B | Thinking-Cap (this model) |
|---|---|---|
| GSM8K (150) | 88.0% @ 533 | 91.3% @ 136 (−74% tokens) |
| HumanEval (164) | 92.7% @ 728 | 92.1% @ 669 (−8% tokens) |
| MBPP (100) | 71.0% @ 846 | 73.0% @ 689 (−19% tokens) |
| average accuracy | 83.9% | 85.5% |
| total tokens | 2107 | 1494 (−29%) |
Adaptive concision. The token cut scales to how compressible each task is — dramatic on math (reasoning is cheap), modest on code.
A frequent failure mode of length-penalized RL is repetitive degeneration ("anxiety") in agentic use. Tested on 24 loop-prone prompts (hard reasoning, eval-framing, ambiguous), greedy decode, with a repeated-span detector:
| Model | Loop rate | Mean tokens | Runaway completions |
|---|---|---|---|
| Base Ornith-9B | 0.0% | 1014 | 3 |
| Thinking-Cap | 0.0% | 607 | 1 |
Neither loops; the tune produces fewer runaway completions than the base model. The correctness reinforcement kept the accuracy up.
GGUF quants preserve the behavior. Benchmarked via Ollama (one harness for both, so the difference isolates the quantization effect; absolute numbers run slightly different from the transformers table above because of the GGUF chat-template path):
| Format | Size | GSM8K | HumanEval | MBPP | avg |
|---|---|---|---|---|---|
| f16 | 17 GB | 91.3% | 90.9% | 77.0% | 86.4% |
| Q8_0 | 9.5 GB | 90.0% | 89.0% | 76.0% | 85.0% |
Q8 lands within 1–3 problems of f16 on every benchmark (−1.4 pts average) with identical token behavior — quantization is effectively free. Q8_0 at half the memory is the recommended daily driver; f16 is the publish-precision reference. (Lower quants are not recommended: a Q4 of the base model failed to complete the eval at all.)
Correctness-gated, length-penalized GRPO (RL with verifiable rewards). Per completion:
reward = correctness # 1.0 if the answer/tests pass, else 0.0
− (λ · normalized_length if correct) # length penalty ONLY when correct
+ format_bonus # well-formed <think>…</think> + answer
The gate is the safety property. The length penalty applies only to correct answers, so the model can never trade accuracy for brevity. GRPO's group-relative advantage means compression only happens on prompts the model has already mastered; a KL leash to the frozen base model and a λ ramp (off → max over ~300 steps) keep it stable and prevent degeneration. The result is self-limiting: the length curve hits a floor and plateaus (checkpoints at steps 400–800 are a tied plateau).
Trained with LoRA (r32/a64, all-linear, MoE router frozen) via Unsloth + TRL GRPOTrainer
on an NVIDIA DGX Spark (GB10). Reward signal reuses the same verifiable checkers used to grade
the model, so training and evaluation share one oracle.
Full recipe, reward harness, and eval scripts: github.com/khudgins/ornith-thinking-cap. Method:
ThinkingCap (correctness-gated length-penalized GRPO) — an independent reproduction of the
ThinkingCap model series by BottleCap AI
(announcement ·
bottlecapai/ThinkingCap-Qwen3.6-27B),
applied here to a smaller coding model. Not affiliated with or endorsed by BottleCap AI. Base model:
deepreinforce-ai/Ornith-1.0-9B (MIT per its HF tag), itself post-trained on Qwen3.5 (Apache
2.0). See License for the full chain. Tooling: Unsloth, TRL, llama.cpp, Ollama.
MIT. This fine-tune — the LoRA adapter and any weights merged from it — is released under the MIT License. It adds no new base weights; it's a low-rank delta trained on top of the base.
Upstream provenance and terms, stated precisely:
deepreinforce-ai/Ornith-1.0-9B
is tagged MIT on Hugging Face; the inherited base weights are governed by that grant.Qwen3_5ForConditionalGeneration).5 commits