High-fidelity 3-bit Qwen3.8 for long-horizon agents.
54 GB β 12.3 GB Β· +0.02% PPL Β· 93.2% Top-1 Agreement Β· 0.031 KLD Β· 262K Context
OrcaRouter AI Gateway Β· X Β· Discord Β· GitHub Β· All Models
27B reasoning. 12.3 GB.
OrcaSAQ2 27B compresses Qwen3.8-27B from a 54 GB BF16 checkpoint to 12.3 GB while preserving extremely high fidelity to the original model.
Built for: long-horizon agents Β· coding Β· tool use Β· reasoning Β· stateful execution
OrcaSAQ2 is a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter and its research team behind.
It is optimized around one goal: Preserve as much useful model behavior as possible inside a practical GPU memory envelope.
The resulting checkpoint provides:
| Metric | BF16 | OrcaSAQ2 |
|---|---|---|
| Checkpoint | 54 GB | 12.3 GB |
| Relative size | 100% | 22.8% |
| Storage reduction | β | 77.2% |
| Decoder precision | 16-bit | 3.21 bpw avg. |
| Perplexity | 5.6468 | 5.6482 |
| PPL delta | β | +0.02% |
| Top-1 agreement | 100% | 93.2% |
| Mean KLD | β | 0.031 |
| Context | 262K | 262K |
4.4Γ smaller. +0.02% perplexity.
The point is not 3-bit.
The point is what survives at 3-bit.
All numbers below are measured using these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.
16,376 predicted tokens
| Build | Size | Decoder Bits | Mean KLD β | Top-1 Agreement β | PPL β |
|---|---|---|---|---|---|
| Qwen3.8-27B BF16 | 54 GB | 16 | β | 100% | 5.6468 |
| OrcaSAQ2 27B | 12.3 GB | 3.21 | 0.031 | 93.2% | 5.6482 |
BF16 5.6468 ββββββββββββββββββββββββββββββββββββββββ
OrcaSAQ2 5.6482 ββββββββββββββββββββββββββββββββββββββββ
Delta: +0.02%
OrcaSAQ2 vs BF16
ββββββββββββββββββββββββββββββββββββββββββββββββββ 93.2%
Qwen3.8-27B BF16
ββββββββββββββββββββββββββββββββββββββββββββββββββ 54.0 GB
OrcaSAQ2
βββββββββββ 12.3 GB
77.2% smaller.
Short benchmarks can hide small degradation.
Agents cannot.
A small model error can change a tool call.
That changes the environment state.
The changed state affects every decision that follows.
Plan
β
Act
β
Observe
β
Decide
β
Recover
β
Repeat
β
...
β
Task Success
Across long trajectories, small errors can compound into large behavioral differences.
That makes long-horizon execution an especially useful stress test for compressed reasoning models.
OrcaSAQ2 performs strongly on long-horizon workloads relative to models in its deployment and parameter class, despite operating from a 12.3 GB checkpoint.
This makes it particularly suitable for:
Perplexity asks:
How similar is the next-token distribution?
Long-horizon evaluation asks:
Can the model still finish the job after many decisions?
For agent models, both matter.
Agent benchmarks depend heavily on the surrounding scaffold, tools, reasoning budget, timeouts and execution environment. The results below are therefore shown as public reference points, not direct apples-to-apples comparisons.
| Model | Reported score |
|---|---|
| Claude Sonnet 4.6 | 79.6 |
| Claude Sonnet 4.5 | 77.2 |
| Gemini 3 | 76.2 |
| OrcaSAQ2 27B | 70.0 |
| Qwen3-Coder-480B-A35B | 69.6 |
| Gemini 2.5 Pro | 63.8 |
| GPT-4.1 | 54.6 |
70.0% SWE-bench Verified from a 12.06 GB 27B checkpoint.
| Model / Agent | Reported score |
|---|---|
| Gemini 3.1 Pro / Terminus 2 | 70.7 |
| Claude Opus 4.6 / Claude Code | 70.1 |
| Claude Opus 4.6 / Terminus 2 | 63.8 |
| Claude Sonnet 4.6 / Claude Code | 58.5 |
| OrcaSAQ2 27B | 58.4 |
| Gemini 3 Flash / Gemini CLI | 56.9 |
| GPT-5.4 / Terminus 2 | 54.8 |
| Claude Sonnet 4.6 / Terminus 2 | 51.5 |
58.4% Terminal-Bench 2.1 while fitting in ~12 GB of checkpoint storage.
Public scores use different agent stacks and should not be interpreted as a strict model-only ranking.
| Base model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForCausalLM |
| Layers | 64 |
| Hidden size | 5120 |
| Hybrid attention | 48 Gated DeltaNet + 16 full-attention layers |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| MTP head | Included |
| Thinking | Supported |
| Tool calling | Supported |
| Checkpoint | 12.3 GB |
| Decoder average | 3.21 bpw |
| Serving | vLLM |
| Vision | Not included |
| License | Apache-2.0 |
Measured under a 15.7 GiB GPU memory cap.
| Configuration | 1 Stream | 8 Streams | 16 Streams | KV Pool |
|---|---|---|---|---|
| vLLM Β· MTP off | 65.3 tok/s | 332 tok/s | 333 tok/s | 29,354 tok |
| vLLM Β· MTP on | 90.1 tok/s | 220 tok/s | 219 tok/s | 14,563 tok |
Single-stream decode
MTP off βββββββββββββββββββββββββββββ 65.3 tok/s
MTP on ββββββββββββββββββββββββββββββββββββββββ
90.1 tok/s
+38% single-stream decode throughput
MTP trades additional compute and KV capacity for stronger interactive decode performance.
It is particularly useful for:
For highly batched workloads, benchmark both configurations.
OrcaSAQ2's checkpoint is 12.3 GB.
That makes deployment possible on hardware that cannot hold the original 54 GB BF16 checkpoint.
16 GB GPU
βββββββββββββββββββββββββββββββββββββββββββββ
β β
β OrcaSAQ2 weights 12.3 GB β
β βββββββββββββββββββββββββββββββββββ β
β β
β Remaining ~3.7 GB β
β ββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββ
Actual usable memory depends on:
A practical starting point for a 16 GB GPU is approximately 32K interactive context, then tune based on the workload.
The model architecture supports up to 262K context.
plan β act β observe β recover β repeat
Repository-scale generation, editing, testing and debugging.
Structured workflows where action-selection quality matters.
Preserving the capabilities of the 27B base model under an aggressive deployment constraint.
A 12.3 GB checkpoint designed around practical inference hardware.
vLLM + MTP + OpenAI-compatible APIs.
One prompt each, first attempt.
The standard SVG test, asked for as an animation.
Chain over the chainring, cranks 180Β° out of phase, parallax background. Pure SMIL, no JavaScript. Used as generated.
Create a html low-poly 3D models of the Statue of Liberty
A single self-contained HTML file: Three.js scene, orbit controls, procedural geometry.
pip install -U vllm huggingface_hub
pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel
hf download orcarouter/OrcaSAQ2-27B \
--local-dir ./OrcaSAQ2-27B
vllm serve ./OrcaSAQ2-27B \
--served-model-name OrcaSAQ2-27B \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed",
)
response = client.chat.completions.create(
model="OrcaSAQ2-27B",
messages=[
{
"role": "user",
"content": "Analyze this repository and plan the next five actions."
}
],
)
print(response.choices[0].message.content)
temperature = 1.0
top_p = 0.95
top_k = 20
Thinking mode is enabled by default.
For agent deployments, benchmark against the actual tool schema, context distribution and reasoning budget used in production.
A low-bit reasoning model should not be judged by checkpoint size alone.
We look at the intersection of:
Footprint Γ BF16 Fidelity Γ Capability Γ Long-Horizon Stability Γ Serving Performance
A useful low-bit model must remain useful after compression.
Perplexity is useful and reproducible.
It is not a complete measure of agentic capability.
Quantization can affect:
reasoning
β
planning
β
tool selection
β
state tracking
β
recovery
β
task completion
That is why OrcaSAQ2 reports BF16 fidelity metrics alongside downstream and long-horizon evaluation.
OrcaSAQ2 uses a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter.
The implementation is optimized to preserve model quality under a strict deployment-memory target.
Detailed quantization methodology, calibration strategy, precision allocation and packing techniques are not currently disclosed.
Open multi-model code review.
Record, replay, fork and debug AI-agent runs.
Self-hosted multi-model AI infrastructure.
Open model. Open harness. Open bill.
@misc{qwen38,
title = {Qwen3.8-Max: A New Bar for Coding and Cowork},
author = {{Qwen Team}},
year = {2026},
month = {August},
url = {https://qwen.ai/blog?id=qwen3.8}
}
Apache-2.0
Inherited from:
Quantization does not change the underlying license obligations.
Route Smarter Β· Ship Safer Β· Spend Less
10 commits
High-fidelity 3-bit Qwen3.8 for long-horizon agents.
54 GB β 12.3 GB Β· +0.02% PPL Β· 93.2% Top-1 Agreement Β· 0.031 KLD Β· 262K Context
OrcaRouter AI Gateway Β· X Β· Discord Β· GitHub Β· All Models
27B reasoning. 12.3 GB.
OrcaSAQ2 27B compresses Qwen3.8-27B from a 54 GB BF16 checkpoint to 12.3 GB while preserving extremely high fidelity to the original model.
Built for: long-horizon agents Β· coding Β· tool use Β· reasoning Β· stateful execution
OrcaSAQ2 is a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter and its research team behind.
It is optimized around one goal: Preserve as much useful model behavior as possible inside a practical GPU memory envelope.
The resulting checkpoint provides:
| Metric | BF16 | OrcaSAQ2 |
|---|---|---|
| Checkpoint | 54 GB | 12.3 GB |
| Relative size | 100% | 22.8% |
| Storage reduction | β | 77.2% |
| Decoder precision | 16-bit | 3.21 bpw avg. |
| Perplexity | 5.6468 | 5.6482 |
| PPL delta | β | +0.02% |
| Top-1 agreement | 100% | 93.2% |
| Mean KLD | β | 0.031 |
| Context | 262K | 262K |
4.4Γ smaller. +0.02% perplexity.
The point is not 3-bit.
The point is what survives at 3-bit.
All numbers below are measured using these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.
16,376 predicted tokens
| Build | Size | Decoder Bits | Mean KLD β | Top-1 Agreement β | PPL β |
|---|---|---|---|---|---|
| Qwen3.8-27B BF16 | 54 GB | 16 | β | 100% | 5.6468 |
| OrcaSAQ2 27B | 12.3 GB | 3.21 | 0.031 | 93.2% | 5.6482 |
BF16 5.6468 ββββββββββββββββββββββββββββββββββββββββ
OrcaSAQ2 5.6482 ββββββββββββββββββββββββββββββββββββββββ
Delta: +0.02%
OrcaSAQ2 vs BF16
ββββββββββββββββββββββββββββββββββββββββββββββββββ 93.2%
Qwen3.8-27B BF16
ββββββββββββββββββββββββββββββββββββββββββββββββββ 54.0 GB
OrcaSAQ2
βββββββββββ 12.3 GB
77.2% smaller.
Short benchmarks can hide small degradation.
Agents cannot.
A small model error can change a tool call.
That changes the environment state.
The changed state affects every decision that follows.
Plan
β
Act
β
Observe
β
Decide
β
Recover
β
Repeat
β
...
β
Task Success
Across long trajectories, small errors can compound into large behavioral differences.
That makes long-horizon execution an especially useful stress test for compressed reasoning models.
OrcaSAQ2 performs strongly on long-horizon workloads relative to models in its deployment and parameter class, despite operating from a 12.3 GB checkpoint.
This makes it particularly suitable for:
Perplexity asks:
How similar is the next-token distribution?
Long-horizon evaluation asks:
Can the model still finish the job after many decisions?
For agent models, both matter.
Agent benchmarks depend heavily on the surrounding scaffold, tools, reasoning budget, timeouts and execution environment. The results below are therefore shown as public reference points, not direct apples-to-apples comparisons.
| Model | Reported score |
|---|---|
| Claude Sonnet 4.6 | 79.6 |
| Claude Sonnet 4.5 | 77.2 |
| Gemini 3 | 76.2 |
| OrcaSAQ2 27B | 70.0 |
| Qwen3-Coder-480B-A35B | 69.6 |
| Gemini 2.5 Pro | 63.8 |
| GPT-4.1 | 54.6 |
70.0% SWE-bench Verified from a 12.06 GB 27B checkpoint.
| Model / Agent | Reported score |
|---|---|
| Gemini 3.1 Pro / Terminus 2 | 70.7 |
| Claude Opus 4.6 / Claude Code | 70.1 |
| Claude Opus 4.6 / Terminus 2 | 63.8 |
| Claude Sonnet 4.6 / Claude Code | 58.5 |
| OrcaSAQ2 27B | 58.4 |
| Gemini 3 Flash / Gemini CLI | 56.9 |
| GPT-5.4 / Terminus 2 | 54.8 |
| Claude Sonnet 4.6 / Terminus 2 | 51.5 |
58.4% Terminal-Bench 2.1 while fitting in ~12 GB of checkpoint storage.
Public scores use different agent stacks and should not be interpreted as a strict model-only ranking.
| Base model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForCausalLM |
| Layers | 64 |
| Hidden size | 5120 |
| Hybrid attention | 48 Gated DeltaNet + 16 full-attention layers |
| Context | 262,144 tokens |
| Vocabulary | 248,320 |
| MTP head | Included |
| Thinking | Supported |
| Tool calling | Supported |
| Checkpoint | 12.3 GB |
| Decoder average | 3.21 bpw |
| Serving | vLLM |
| Vision | Not included |
| License | Apache-2.0 |
Measured under a 15.7 GiB GPU memory cap.
| Configuration | 1 Stream | 8 Streams | 16 Streams | KV Pool |
|---|---|---|---|---|
| vLLM Β· MTP off | 65.3 tok/s | 332 tok/s | 333 tok/s | 29,354 tok |
| vLLM Β· MTP on | 90.1 tok/s | 220 tok/s | 219 tok/s | 14,563 tok |
Single-stream decode
MTP off βββββββββββββββββββββββββββββ 65.3 tok/s
MTP on ββββββββββββββββββββββββββββββββββββββββ
90.1 tok/s
+38% single-stream decode throughput
MTP trades additional compute and KV capacity for stronger interactive decode performance.
It is particularly useful for:
For highly batched workloads, benchmark both configurations.
OrcaSAQ2's checkpoint is 12.3 GB.
That makes deployment possible on hardware that cannot hold the original 54 GB BF16 checkpoint.
16 GB GPU
βββββββββββββββββββββββββββββββββββββββββββββ
β β
β OrcaSAQ2 weights 12.3 GB β
β βββββββββββββββββββββββββββββββββββ β
β β
β Remaining ~3.7 GB β
β ββββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββ
Actual usable memory depends on:
A practical starting point for a 16 GB GPU is approximately 32K interactive context, then tune based on the workload.
The model architecture supports up to 262K context.
plan β act β observe β recover β repeat
Repository-scale generation, editing, testing and debugging.
Structured workflows where action-selection quality matters.
Preserving the capabilities of the 27B base model under an aggressive deployment constraint.
A 12.3 GB checkpoint designed around practical inference hardware.
vLLM + MTP + OpenAI-compatible APIs.
One prompt each, first attempt.
The standard SVG test, asked for as an animation.
Chain over the chainring, cranks 180Β° out of phase, parallax background. Pure SMIL, no JavaScript. Used as generated.
Create a html low-poly 3D models of the Statue of Liberty
A single self-contained HTML file: Three.js scene, orbit controls, procedural geometry.
pip install -U vllm huggingface_hub
pip install git+https://github.com/Continuum-AI-Corp/OrcaSAQ2-kernel
hf download orcarouter/OrcaSAQ2-27B \
--local-dir ./OrcaSAQ2-27B
vllm serve ./OrcaSAQ2-27B \
--served-model-name OrcaSAQ2-27B \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed",
)
response = client.chat.completions.create(
model="OrcaSAQ2-27B",
messages=[
{
"role": "user",
"content": "Analyze this repository and plan the next five actions."
}
],
)
print(response.choices[0].message.content)
temperature = 1.0
top_p = 0.95
top_k = 20
Thinking mode is enabled by default.
For agent deployments, benchmark against the actual tool schema, context distribution and reasoning budget used in production.
A low-bit reasoning model should not be judged by checkpoint size alone.
We look at the intersection of:
Footprint Γ BF16 Fidelity Γ Capability Γ Long-Horizon Stability Γ Serving Performance
A useful low-bit model must remain useful after compression.
Perplexity is useful and reproducible.
It is not a complete measure of agentic capability.
Quantization can affect:
reasoning
β
planning
β
tool selection
β
state tracking
β
recovery
β
task completion
That is why OrcaSAQ2 reports BF16 fidelity metrics alongside downstream and long-horizon evaluation.
OrcaSAQ2 uses a proprietary sensitivity-aware mixed-precision quantization system developed by OrcaRouter.
The implementation is optimized to preserve model quality under a strict deployment-memory target.
Detailed quantization methodology, calibration strategy, precision allocation and packing techniques are not currently disclosed.
Open multi-model code review.
Record, replay, fork and debug AI-agent runs.
Self-hosted multi-model AI infrastructure.
Open model. Open harness. Open bill.
@misc{qwen38,
title = {Qwen3.8-Max: A New Bar for Coding and Cowork},
author = {{Qwen Team}},
year = {2026},
month = {August},
url = {https://qwen.ai/blog?id=qwen3.8}
}
Apache-2.0
Inherited from:
Quantization does not change the underlying license obligations.
Route Smarter Β· Ship Safer Β· Spend Less
10 commits