RouteWeaver: experimental cache-safe speculative inference and CPU/GPU transfer optimization for Qwen 27B on low-VRAM NVIDIA GPUs.
Python
31
10 commits
updated Oct 6, 2026

Local inference, further.
Experimental inference-engine optimizations for running large language models on memory-constrained hardware. Built on llama.cpp: cache-safe speculative decoding, queued CPU-to-GPU weight transfers, selective weight placement, and focused CPU/CUDA kernels.
Run Q3 · Results · Hardware · How it works · Contribute
The October 5 configuration improved generation throughput by 14.5–15.4% over the previous same-GGUF setup across three short fixtures, five seeds each. No change to the weights, allocated context, Q4_0 KV, draft depth or sampling between the paired arms.
Mean generated tok/s; parentheses show mean accepted/drafted ratio:
| Fixture | Previous setup | Cache-safe setup | Change |
|---|---|---|---|
| Rust coding | 22.94 (77.1%) | 26.36 (77.1%) | +14.9% |
| C++ coding | 20.97 (68.3%) | 24.01 (68.2%) | +14.5% |
| Reasoning | 15.90 (44.8%) | 18.34 (45.1%) | +15.4% |
All 15 paired output messages matched exactly. These are throughput fixtures, not coding-quality scores or a promise that every conversation reaches these speeds. The measured improvement combines D-CFR, placement and buffer changes; it is not an isolated D-CFR-only speedup.
A separate synthetic stress test processed 121,821 input tokens in a 131,072-token window and then generated at 19.14 tok/s. It demonstrates memory fit and operation near a full context, not long-context reasoning accuracy. The lowest sampled free VRAM was 366 MiB; other GPU applications can exhaust that margin.
Full ranges, acceptance, occupied-context tests and limitations · Machine-readable measurements
The current Q3 release targets the hash-pinned RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf, not every Qwen quant.
For Linux x86-64 with AVX2/FMA and one NVIDIA GPU with at least 12 GB VRAM. Install the build prerequisites first; the setup command detects a compatible installed CUDA/compiler pair.
git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./routeweaver setup --context 98304
./routeweaver start --context 98304
Use --context 131072 on both commands for 128K. Setup verifies/downloads the
model, builds locally, tests a bounded set of settings, and saves a validated
profile. It keeps your context, quant and Q4_0 KV fixed. Existing GPU compute
processes must be stopped by their owner; setup never closes them automatically.
The default tuning budget is 15 minutes, excluding download/build and model-free checks. A failed or inconclusive tune does not replace an existing profile. This is the best passing configuration sampled, not a guaranteed global optimum. A real 96K tuning session passed on the RTX 3060 development machine; other hardware remains experimental. See validation scope and full instructions.
Prefer explicit settings? This remains a fully independent installation path:
git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./scripts/download-q3.sh
# Point these at a CUDA installation and supported host compilers.
CUDA_TOOLKIT_ROOT=/usr/local/cuda CC=gcc-15 CXX=g++-15 \
./scripts/build-q3.sh
./scripts/check-q3.sh
# 96K default; use CONTEXT=131072 for the 128K preset.
./scripts/serve-q3.sh
Tested hardware: RTX 3060 12 GB, Core i5-12400, 16 GB RAM, Fedora Linux. Stop other GPU model servers first. Read the complete setup, vision, client connection and rollback guide. The server binds to localhost and exposes an OpenAI-compatible API.
The fresh release builds these components together from pinned source. Historical measurements used separately built runtime libraries; toolchain and packaging details are disclosed in the results. This repository is an experimental fork/patch set, not an upstream llama.cpp release.
| Status | Scope |
|---|---|
| Locally tested | RTX 3060 12 GB + i5-12400, this Q3 model, 96K/128K |
| Community reports | Older configurations; separate from validation of this release |
| Experimental | Other NVIDIA GPUs, CPUs, VRAM sizes and operating systems |
| Not validated | AMD/Intel/Apple GPU backends, multi-GPU, other model architectures |
Expect to rebuild and retune for another machine. D-CFR is architecture-specific; prefetch gains depend on how much work crosses the CPU/GPU boundary. Compatibility, starting tests and report requirements
The repository URL stays unchanged so existing links, issues and history keep working. RouteWeaver is the project brand; D-CFR names the recurrent-state technique.
Independent reproduction matters. Share matched prompts and seeds, model hashes, draft acceptance, occupied context, memory and timings. Report slower results too. See CONTRIBUTING or open a hardware report.
Code is MIT licensed. Upstream and model credits are in NOTICE.
RouteWeaver: experimental cache-safe speculative inference and CPU/GPU transfer optimization for Qwen 27B on low-VRAM NVIDIA GPUs.
Python
31
10 commits
updated Oct 6, 2026

Local inference, further.
Experimental inference-engine optimizations for running large language models on memory-constrained hardware. Built on llama.cpp: cache-safe speculative decoding, queued CPU-to-GPU weight transfers, selective weight placement, and focused CPU/CUDA kernels.
Run Q3 · Results · Hardware · How it works · Contribute
The October 5 configuration improved generation throughput by 14.5–15.4% over the previous same-GGUF setup across three short fixtures, five seeds each. No change to the weights, allocated context, Q4_0 KV, draft depth or sampling between the paired arms.
Mean generated tok/s; parentheses show mean accepted/drafted ratio:
| Fixture | Previous setup | Cache-safe setup | Change |
|---|---|---|---|
| Rust coding | 22.94 (77.1%) | 26.36 (77.1%) | +14.9% |
| C++ coding | 20.97 (68.3%) | 24.01 (68.2%) | +14.5% |
| Reasoning | 15.90 (44.8%) | 18.34 (45.1%) | +15.4% |
All 15 paired output messages matched exactly. These are throughput fixtures, not coding-quality scores or a promise that every conversation reaches these speeds. The measured improvement combines D-CFR, placement and buffer changes; it is not an isolated D-CFR-only speedup.
A separate synthetic stress test processed 121,821 input tokens in a 131,072-token window and then generated at 19.14 tok/s. It demonstrates memory fit and operation near a full context, not long-context reasoning accuracy. The lowest sampled free VRAM was 366 MiB; other GPU applications can exhaust that margin.
Full ranges, acceptance, occupied-context tests and limitations · Machine-readable measurements
The current Q3 release targets the hash-pinned RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf, not every Qwen quant.
For Linux x86-64 with AVX2/FMA and one NVIDIA GPU with at least 12 GB VRAM. Install the build prerequisites first; the setup command detects a compatible installed CUDA/compiler pair.
git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./routeweaver setup --context 98304
./routeweaver start --context 98304
Use --context 131072 on both commands for 128K. Setup verifies/downloads the
model, builds locally, tests a bounded set of settings, and saves a validated
profile. It keeps your context, quant and Q4_0 KV fixed. Existing GPU compute
processes must be stopped by their owner; setup never closes them automatically.
The default tuning budget is 15 minutes, excluding download/build and model-free checks. A failed or inconclusive tune does not replace an existing profile. This is the best passing configuration sampled, not a guaranteed global optimum. A real 96K tuning session passed on the RTX 3060 development machine; other hardware remains experimental. See validation scope and full instructions.
Prefer explicit settings? This remains a fully independent installation path:
git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./scripts/download-q3.sh
# Point these at a CUDA installation and supported host compilers.
CUDA_TOOLKIT_ROOT=/usr/local/cuda CC=gcc-15 CXX=g++-15 \
./scripts/build-q3.sh
./scripts/check-q3.sh
# 96K default; use CONTEXT=131072 for the 128K preset.
./scripts/serve-q3.sh
Tested hardware: RTX 3060 12 GB, Core i5-12400, 16 GB RAM, Fedora Linux. Stop other GPU model servers first. Read the complete setup, vision, client connection and rollback guide. The server binds to localhost and exposes an OpenAI-compatible API.
The fresh release builds these components together from pinned source. Historical measurements used separately built runtime libraries; toolchain and packaging details are disclosed in the results. This repository is an experimental fork/patch set, not an upstream llama.cpp release.
| Status | Scope |
|---|---|
| Locally tested | RTX 3060 12 GB + i5-12400, this Q3 model, 96K/128K |
| Community reports | Older configurations; separate from validation of this release |
| Experimental | Other NVIDIA GPUs, CPUs, VRAM sizes and operating systems |
| Not validated | AMD/Intel/Apple GPU backends, multi-GPU, other model architectures |
Expect to rebuild and retune for another machine. D-CFR is architecture-specific; prefetch gains depend on how much work crosses the CPU/GPU boundary. Compatibility, starting tests and report requirements
The repository URL stays unchanged so existing links, issues and history keep working. RouteWeaver is the project brand; D-CFR names the recurrent-state technique.
Independent reproduction matters. Share matched prompts and seeds, model hashes, draft acceptance, occupied context, memory and timings. Report slower results too. See CONTRIBUTING or open a hardware report.
Code is MIT licensed. Upstream and model credits are in NOTICE.