kadenball/qwen38-27b-rtx3060-dcfr

RouteWeaver: experimental cache-safe speculative inference and CPU/GPU transfer optimization for Qwen 27B on low-VRAM NVIDIA GPUs.

Python

31

10 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Routeweaver - Serve 27b fast on low vram set ups (r/LocalLLM)

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put…

1

Oct 6, 2026

Routeweaver - Serve 27b fast on low vram set ups (r/LocalLLaMA)

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put…

1

Oct 6, 2026

README

RouteWeaver

RouteWeaver

Local inference, further.

Experimental inference-engine optimizations for running large language models on memory-constrained hardware. Built on llama.cpp: cache-safe speculative decoding, queued CPU-to-GPU weight transfers, selective weight placement, and focused CPU/CUDA kernels.

Run Q3 · Results · Hardware · How it works · Contribute

Latest: Qwen 27B Q3, 128K context, one RTX 3060 12 GB

The October 5 configuration improved generation throughput by 14.5–15.4% over the previous same-GGUF setup across three short fixtures, five seeds each. No change to the weights, allocated context, Q4_0 KV, draft depth or sampling between the paired arms.

Mean generated tok/s; parentheses show mean accepted/drafted ratio:

FixturePrevious setupCache-safe setupChange
Rust coding22.94 (77.1%)26.36 (77.1%)+14.9%
C++ coding20.97 (68.3%)24.01 (68.2%)+14.5%
Reasoning15.90 (44.8%)18.34 (45.1%)+15.4%

All 15 paired output messages matched exactly. These are throughput fixtures, not coding-quality scores or a promise that every conversation reaches these speeds. The measured improvement combines D-CFR, placement and buffer changes; it is not an isolated D-CFR-only speedup.

A separate synthetic stress test processed 121,821 input tokens in a 131,072-token window and then generated at 19.14 tok/s. It demonstrates memory fit and operation near a full context, not long-context reasoning accuracy. The lowest sampled free VRAM was 366 MiB; other GPU applications can exhaust that margin.

Full ranges, acceptance, occupied-context tests and limitations · Machine-readable measurements

Start here

The current Q3 release targets the hash-pinned RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf, not every Qwen quant.

Guided setup and automatic tuning (experimental)

For Linux x86-64 with AVX2/FMA and one NVIDIA GPU with at least 12 GB VRAM. Install the build prerequisites first; the setup command detects a compatible installed CUDA/compiler pair.

git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./routeweaver setup --context 98304
./routeweaver start --context 98304

Use --context 131072 on both commands for 128K. Setup verifies/downloads the model, builds locally, tests a bounded set of settings, and saves a validated profile. It keeps your context, quant and Q4_0 KV fixed. Existing GPU compute processes must be stopped by their owner; setup never closes them automatically.

The default tuning budget is 15 minutes, excluding download/build and model-free checks. A failed or inconclusive tune does not replace an existing profile. This is the best passing configuration sampled, not a guaranteed global optimum. A real 96K tuning session passed on the RTX 3060 development machine; other hardware remains experimental. See validation scope and full instructions.

Manual installation

Prefer explicit settings? This remains a fully independent installation path:

git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./scripts/download-q3.sh

# Point these at a CUDA installation and supported host compilers.
CUDA_TOOLKIT_ROOT=/usr/local/cuda CC=gcc-15 CXX=g++-15 \
  ./scripts/build-q3.sh
./scripts/check-q3.sh

# 96K default; use CONTEXT=131072 for the 128K preset.
./scripts/serve-q3.sh

Tested hardware: RTX 3060 12 GB, Core i5-12400, 16 GB RAM, Fedora Linux. Stop other GPU model servers first. Read the complete setup, vision, client connection and rollback guide. The server binds to localhost and exposes an OpenAI-compatible API.

Where the gain comes from

  • Cache-safe D-CFR: preserve the recurrent base, pending updates and rollback metadata together, making compact speculative state compatible with prompt reuse.
  • More GPU-resident weights: spend the recovered VRAM on eight additional blocks' gate/up weights at the same context allocation.
  • Queued transfers: three 48 MiB staging regions overlap eligible host-weight transfers with execution and avoid a redundant device copy.
  • Specialized kernels: reuse IQ3 decoding across small CPU query batches and reduce Q4_0 attention workspace on the measured CUDA path.

The fresh release builds these components together from pinned source. Historical measurements used separately built runtime libraries; toolchain and packaging details are disclosed in the results. This repository is an experimental fork/patch set, not an upstream llama.cpp release.

Hardware support is evidence-based

StatusScope
Locally testedRTX 3060 12 GB + i5-12400, this Q3 model, 96K/128K
Community reportsOlder configurations; separate from validation of this release
ExperimentalOther NVIDIA GPUs, CPUs, VRAM sizes and operating systems
Not validatedAMD/Intel/Apple GPU backends, multi-GPU, other model architectures

Expect to rebuild and retune for another machine. D-CFR is architecture-specific; prefetch gains depend on how much work crosses the CPU/GPU boundary. Compatibility, starting tests and report requirements

Project map

The repository URL stays unchanged so existing links, issues and history keep working. RouteWeaver is the project brand; D-CFR names the recurrent-state technique.

Contribute

Independent reproduction matters. Share matched prompts and seeds, model hashes, draft acceptance, occupied context, memory and timings. Report slower results too. See CONTRIBUTING or open a hardware report.

Code is MIT licensed. Upstream and model credits are in NOTICE.

benchmarks
cuda
llama-cpp
local-llm
qwen
routeweaver
rtx-3060
speculative-decoding

kadenball/qwen38-27b-rtx3060-dcfr

RouteWeaver: experimental cache-safe speculative inference and CPU/GPU transfer optimization for Qwen 27B on low-VRAM NVIDIA GPUs.

Python

31

10 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Routeweaver - Serve 27b fast on low vram set ups (r/LocalLLM)

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put…

1

Oct 6, 2026

Routeweaver - Serve 27b fast on low vram set ups (r/LocalLLaMA)

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put…

1

Oct 6, 2026

README

RouteWeaver

RouteWeaver

Local inference, further.

Experimental inference-engine optimizations for running large language models on memory-constrained hardware. Built on llama.cpp: cache-safe speculative decoding, queued CPU-to-GPU weight transfers, selective weight placement, and focused CPU/CUDA kernels.

Run Q3 · Results · Hardware · How it works · Contribute

Latest: Qwen 27B Q3, 128K context, one RTX 3060 12 GB

The October 5 configuration improved generation throughput by 14.5–15.4% over the previous same-GGUF setup across three short fixtures, five seeds each. No change to the weights, allocated context, Q4_0 KV, draft depth or sampling between the paired arms.

Mean generated tok/s; parentheses show mean accepted/drafted ratio:

FixturePrevious setupCache-safe setupChange
Rust coding22.94 (77.1%)26.36 (77.1%)+14.9%
C++ coding20.97 (68.3%)24.01 (68.2%)+14.5%
Reasoning15.90 (44.8%)18.34 (45.1%)+15.4%

All 15 paired output messages matched exactly. These are throughput fixtures, not coding-quality scores or a promise that every conversation reaches these speeds. The measured improvement combines D-CFR, placement and buffer changes; it is not an isolated D-CFR-only speedup.

A separate synthetic stress test processed 121,821 input tokens in a 131,072-token window and then generated at 19.14 tok/s. It demonstrates memory fit and operation near a full context, not long-context reasoning accuracy. The lowest sampled free VRAM was 366 MiB; other GPU applications can exhaust that margin.

Full ranges, acceptance, occupied-context tests and limitations · Machine-readable measurements

Start here

The current Q3 release targets the hash-pinned RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_XXS-mtp.gguf, not every Qwen quant.

Guided setup and automatic tuning (experimental)

For Linux x86-64 with AVX2/FMA and one NVIDIA GPU with at least 12 GB VRAM. Install the build prerequisites first; the setup command detects a compatible installed CUDA/compiler pair.

git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./routeweaver setup --context 98304
./routeweaver start --context 98304

Use --context 131072 on both commands for 128K. Setup verifies/downloads the model, builds locally, tests a bounded set of settings, and saves a validated profile. It keeps your context, quant and Q4_0 KV fixed. Existing GPU compute processes must be stopped by their owner; setup never closes them automatically.

The default tuning budget is 15 minutes, excluding download/build and model-free checks. A failed or inconclusive tune does not replace an existing profile. This is the best passing configuration sampled, not a guaranteed global optimum. A real 96K tuning session passed on the RTX 3060 development machine; other hardware remains experimental. See validation scope and full instructions.

Manual installation

Prefer explicit settings? This remains a fully independent installation path:

git clone https://github.com/kadenball/qwen38-27b-rtx3060-dcfr.git
cd qwen38-27b-rtx3060-dcfr
./scripts/download-q3.sh

# Point these at a CUDA installation and supported host compilers.
CUDA_TOOLKIT_ROOT=/usr/local/cuda CC=gcc-15 CXX=g++-15 \
  ./scripts/build-q3.sh
./scripts/check-q3.sh

# 96K default; use CONTEXT=131072 for the 128K preset.
./scripts/serve-q3.sh

Tested hardware: RTX 3060 12 GB, Core i5-12400, 16 GB RAM, Fedora Linux. Stop other GPU model servers first. Read the complete setup, vision, client connection and rollback guide. The server binds to localhost and exposes an OpenAI-compatible API.

Where the gain comes from

  • Cache-safe D-CFR: preserve the recurrent base, pending updates and rollback metadata together, making compact speculative state compatible with prompt reuse.
  • More GPU-resident weights: spend the recovered VRAM on eight additional blocks' gate/up weights at the same context allocation.
  • Queued transfers: three 48 MiB staging regions overlap eligible host-weight transfers with execution and avoid a redundant device copy.
  • Specialized kernels: reuse IQ3 decoding across small CPU query batches and reduce Q4_0 attention workspace on the measured CUDA path.

The fresh release builds these components together from pinned source. Historical measurements used separately built runtime libraries; toolchain and packaging details are disclosed in the results. This repository is an experimental fork/patch set, not an upstream llama.cpp release.

Hardware support is evidence-based

StatusScope
Locally testedRTX 3060 12 GB + i5-12400, this Q3 model, 96K/128K
Community reportsOlder configurations; separate from validation of this release
ExperimentalOther NVIDIA GPUs, CPUs, VRAM sizes and operating systems
Not validatedAMD/Intel/Apple GPU backends, multi-GPU, other model architectures

Expect to rebuild and retune for another machine. D-CFR is architecture-specific; prefetch gains depend on how much work crosses the CPU/GPU boundary. Compatibility, starting tests and report requirements

Project map

The repository URL stays unchanged so existing links, issues and history keep working. RouteWeaver is the project brand; D-CFR names the recurrent-state technique.

Contribute

Independent reproduction matters. Share matched prompts and seeds, model hashes, draft acceptance, occupied context, memory and timings. Report slower results too. See CONTRIBUTING or open a hardware report.

Code is MIT licensed. Upstream and model credits are in NOTICE.

benchmarks
cuda
llama-cpp
local-llm
qwen
routeweaver
rtx-3060
speculative-decoding