bodhi37/Qwen3.8-27B-12GBVRAM-Recipe

Qwen3.8-27B on 12GB VRAM + 8GB RAM: 64k context, 16.8 tok/s decode, stock llama.cpp pinned build

Shell

0

1 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec (r/LocalLLaMA)

I got Qwen3.8-27B running at \~18 tok/sec decode & \~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3\_S quant (\~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3\_S quant, achieving similar speeds…

2

Oct 6, 2026

README

Qwen3.8-27B on 12 GB VRAM + 8 GB RAM

Runs on one 12 GB card. Stock llama.cpp at 254b177. No fork, no patches.

Result:

metricmeasured
Decode16.8 tok/s median over 163 requests (p90 18.0, max 18.5). 18.1-18.4 on a 30-token prompt, 17.0 on 4-11k, 13.5 at 57k
Prefill280 tok/s median over the same 163 (max 673). 605 on a 4-11k prompt, 533 on a 57k prompt
Context64000 tokens. A 57232-token prompt prefills at 533 tok/s, no OOM. One slot
VRAM11724/12282 MiB used idle (230 free). 11947 used peak on the 57k request (7 free)
RAM2.1-2.2 GiB idle, 2.8 GiB after one 57k request, 4.0 GiB after three 11k requests. Budget 8 GB
Load1.8-2.0 s to model loaded when the 11 GiB file is in page cache, about 10 s cold

This is not BF16. Base is 53.8 GB at 16-bit and does not fit. This runs the GSQ-RCO IQ3_S GGUF at 3.5 bits/weight. Authors report 99.8% of BF16 on task average. Details and caveats under Quality.


Models

Same flags, same budget for both. Pick one:

buildrepofilesize
StockISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUFQwen3.8-27B-GSQ-RCO-IQ3_S.gguf11.0 GiB
Abliterated0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUFRVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S.gguf11.0 GiB
Abliterated + MTP head (what this box runs)same repoRVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf11.40 GiB

Measured on the same card, same protocol: 3 requests of 10801 tokens, 256-token replies. These are 11k prompts, so VRAM peaks below the 57k figure above:

metricStock IQ3_SAbliterated -mtp
Prefill600 tok/s564 tok/s (-6%)
Decode16.8 tok/s15.6 tok/s (-7%)
VRAM idle to peak11724 to 11932 MiB11758 to 11758 MiB
Free at peak22 MiB196 MiB
RSS peak3.9 GiB4.3 GiB

The abliterated file is the ISTA quant with weights edited in place. All 851 tensors match stock except 48 ssm_alpha tensors at F32 instead of BF16 (+22.5 MiB), plus one extra block blk.64 (draft head, 15 tensors, 430 MiB). At -ngl 58 that leaves 7 blocks on the host instead of 6. That accounts for most of the 7% slowdown: moving block 58 to CPU on stock (-ngl 57) costs the same, 16.7 vs 17.7 tok/s. Idle cost is +34 MiB, so the budget below covers both.


Hardware

Reference box: Ryzen 9 9900X 12C/24T, RTX 4070 SUPER 12 GB, 30 GB RAM (8 GB is enough), NVMe, Linux, GCC 16.2.1, CUDA 13.4.

Throughput follows GPU memory bandwidth (plus CPU bandwidth for the host-resident layers). Placement comes from the architecture and quant, so it carries over to other 12 GB cards.


Why it fits

1. Only 16 of 64 layers need KV cache

Qwen3.8-27B is hybrid SSM + attention (general.architecture = qwen35):

layer typecountKV cache
SSM / linear attention48no, recurrent state only
Full attention (3, 7, 11, ... 63)16yes, every 4th layer

An all-attention model with these dims would need about 16 GB of f16 KV at 64k. This one needs 1.1 GB:

16 layers x 4 KV heads x 256 head_dim x 2 (K+V) = 32768 elements/token
q4_0 = 18 bytes / 32 elements                   = 18432 bytes/token
x 64000 tokens                                  = 1125 MiB
KV typeat 64kfits
f164000 MiBno
q8_02125 MiBno
q4_01125 MiByes

-ctk q4_0 -ctv q4_0 is what makes 64k load.

2. 11 GiB of weights, split over PCIe

IQ3_S averages 3.5 bits/weight. File is 11771546784 bytes (11.0 GiB). Sizes from the GGUF header:

tensor groupsizewhere
Blocks 0-579072 MiBGPU (-ngl 58)
output.weight + output_norm682 MiBGPU
Blocks 58-631073 MiBCPU
token_embd.weight388 MiBCPU (-ot token_embd=CPU)

Embedding lookup runs once per request and stays mmapped, so offloading it saves 388 MiB for no measurable cost.

3. Total budget

componentMiB
Weights on GPU (blocks 0-57 + output)9754
KV cache (q4_0, full 64k)1125
CUDA context + graph (idle residual)~845
Measured idle11724
Measured peak (first-use workspace + prompt cache)11947
Free at peak7

Two notes on that last row. 12282 - used overstates free; the driver holds about 330 MiB outside it. Peak is what matters. Observed peaks were 11916-11947 MiB (7-38 MiB free) depending on request. It never OOMed here, but there is no room left for larger -ub, a second slot, or longer context.

  • -ngl 64 OOMs. Blocks 58-63 add 1073 MiB: 10827 + 1125 KV = 11952 MiB before context, graphs, and driver overhead, about 500 MiB over. It fails on KV alloc: allocating 1125.00 MiB on device 0: cudaMalloc failed: out of memory then failed to allocate buffer for kv cache.
  • -ot token_embd=CPU is required to start at all.

RAM

CPU blocks plus embedding table are 1461 MiB of weights, mmapped. RSS: 2.1-2.2 GiB idle, 2.8 GiB after one 57k request, 4.0 GiB after three 10801-token requests (3.9-4.3 GiB across builds). Grows about 0.6 GiB per cached prompt. That is prompt cache, not weights. Budget 8 GB for a server holding several long contexts. At idle smaps_rollup shows 1598 MiB file-backed plus 525 MiB anonymous.

Load shows a brief VmRSS spike to about 11.4 GiB for a second while the GGUF is mapped, then it drops. At that point Pss_File is 11.1 GiB vs Pss_Anon 0.24 GiB, and system MemAvailable falls only about 2.1 GiB. These are clean page-cache pages. 8 GB RAM loads it.


Quality

BF16 is 53.8 GB. On 12 GB VRAM you quantize regardless. Question is which quant.

GSQ-RCO (GSQ, RCO, ISTA DASLab) is non-uniform: each tensor gets its own quant type from a gradient search under a size budget. Sensitive tensors keep precision, redundant ones lose it.

Published numbers, IQ3_S vs BF16 (sizes decimal GB; the 11.8 GB row is the 11.0 GiB file used here):

variantbpwsizeWiki pplZS avgRecoveryAIME25GPQA-DLCB v6
BF1616.0053.8 GB7.0574.34100.0%100.0089.9085.71
GSQ-RCO IQ3_S (this recipe)3.5011.8 GB7.0774.47100.2%100.0089.3985.71
UD-IQ3_S (uniform ref)3.5212.0 GB7.1675.49101.5%96.6789.9084.00

Task average 91.70 vs 91.87. Ties AIME25 and LiveCodeBench, -0.51 on GPQA-Diamond.

Caveats:

  • These are the release authors' benchmarks. Little independent verification.
  • Still 3.5-bit, not original. Perplexity is worse: wiki 7.07 vs 7.05 (+0.3%), C4 11.76 vs 11.45 (+2.7%), FineWeb-Edu 8.34 vs 8.14 (+2.5%). Exact-match ties can hide loss in recall, style, and long-context behavior.
  • Best option at this size anyway. For about 11 GiB you get near-BF16 (GSQ-RCO) or a clearly worse uniform quant. No BF16-shaped option exists on 12 GB.

Reproduce

1. Model

# Stock, IQ3_S (11.0 GiB)
hf download ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF \
  Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf --local-dir ~/models/qwen3.8-27b

# Abliterated, same quant (11.0 GiB; -mtp twin 11.40 GiB)
hf download 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF \
  RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S.gguf --local-dir ~/models/qwen3.8-27b

Name the file exactly. --include "*IQ3_S*" also pulls the -mtp twin. Other sizes in the ISTA repo: IQ3_XXS (10.1 GB, smaller/faster), IQ2_XS / IQ2_S (8.4 and 9.3 GB, clear quality cost). mmproj adds vision; this recipe is text-only.

Both builds use the same flags below. Any other qwen35 Qwen3.8-27B GGUF up to about 11 GiB is a drop-in.

2. Build

./build.sh        # pins commit 254b177, flags in the script

That commit is clean upstream. Pin matters: buffer sizes and --fit change between releases, so retune -ngl if you move.

3. Serve

./serve.sh                                       # defaults, or:
MODEL=/path/to/model.gguf PORT=8105 ./serve.sh

Verbatim:

llama-server \
  -m  ~/models/qwen3.8-27b/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf \
  --host 127.0.0.1 --port 8105 \
  --parallel 1 -c 64000 -b 2048 -ub 96 -t 12 \
  -ctk q4_0 -ctv q4_0 \
  -ot token_embd=CPU \
  -ngl 58

Abliterated differs only in -m:

  -m  ~/models/qwen3.8-27b/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
flagreason
-ngl 5858 of 64 blocks on GPU = 9072 MiB. 64 OOMs. Drop by one on smaller cards
-ot token_embd=CPU388 MiB embedding table off GPU
-ctk q4_0 -ctv q4_0KV 1125 MiB instead of 4000. Needed for 64k. Third-party A/Bs on other models show about +0.4% perplexity; not remeasured here
-c 64000Sized to that KV budget
--parallel 1One slot. KV allocated once; extra clients queue instead of using more memory
-b 2048 -ub 96Physical batch 96 keeps the graph inside the ~845 MiB residual. Higher is faster prefill, less headroom
-t 12Threads for host layers. Half the physical cores; SMT siblings measured worse
--host 127.0.0.1Loopback only. Add --api-key and set CORS if you expose it

Flash attention stays at default auto, which is on with quantized V cache. -fa off -ctv q4_0 fails with quantized V cache requires flash_attn to be enabled. No -fa flag is needed. CUDA graphs are compiled in (GGML_CUDA_GRAPHS=ON, upstream default OFF). The server logs graphs reused per request when active.


What did not fit

  • MTP speculative decoding (--spec-type draft-mtp). Fails two ways. At -c 64000 it never starts: failed to allocate CUDA0 buffer of size 387146240. At -c 32000 it loads then OOMs on the first request (peak 11862 MiB, 92 free). The draft wants its own context, KV, and graph buffers on top of a full allocation. Carrying the -mtp head without --spec-type is fine (+34 MiB).
  • f16 KV. 4000 MiB at 64k. Does not fit.
  • q8_0 KV. 2125 MiB at 64k, i.e. +1000 MiB over q4_0 against single-digit free. Only works around -c 32000 (KV 1063 MiB) or lower -ngl.

Measuring it yourself

  • Decode/prefill: print_timing lines in the server log, e.g. eval time = 47637.25 ms / 809 tokens (58.96 ms per token, 16.96 tokens per second). Decode depends on context: 18.1-18.4 tok/s on a 30-token prompt, 17.0 on 4-11k, 16.8 across 163 requests, 13.5 at 57k.
  • VRAM: sample peak during a request, not idle: nvidia-smi --query-gpu=memory.used,memory.free --format=csv once per second.
  • Tensor sizes: read the GGUF header. 851 tensors with type and dims. Bytes = elements x bits/8 / block_size; consecutive data offsets give the same answer. Keys: qwen35.block_count, qwen35.attention.head_count_kv, qwen35.attention.key_length, qwen35.full_attention_interval.

License

MIT for these scripts, see LICENSE. Models are Apache-2.0 per Hugging Face metadata.

bodhi37/Qwen3.8-27B-12GBVRAM-Recipe

Qwen3.8-27B on 12GB VRAM + 8GB RAM: 64k context, 16.8 tok/s decode, stock llama.cpp pinned build

Shell

0

1 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec (r/LocalLLaMA)

I got Qwen3.8-27B running at \~18 tok/sec decode & \~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3\_S quant (\~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3\_S quant, achieving similar speeds…

2

Oct 6, 2026

README

Qwen3.8-27B on 12 GB VRAM + 8 GB RAM

Runs on one 12 GB card. Stock llama.cpp at 254b177. No fork, no patches.

Result:

metricmeasured
Decode16.8 tok/s median over 163 requests (p90 18.0, max 18.5). 18.1-18.4 on a 30-token prompt, 17.0 on 4-11k, 13.5 at 57k
Prefill280 tok/s median over the same 163 (max 673). 605 on a 4-11k prompt, 533 on a 57k prompt
Context64000 tokens. A 57232-token prompt prefills at 533 tok/s, no OOM. One slot
VRAM11724/12282 MiB used idle (230 free). 11947 used peak on the 57k request (7 free)
RAM2.1-2.2 GiB idle, 2.8 GiB after one 57k request, 4.0 GiB after three 11k requests. Budget 8 GB
Load1.8-2.0 s to model loaded when the 11 GiB file is in page cache, about 10 s cold

This is not BF16. Base is 53.8 GB at 16-bit and does not fit. This runs the GSQ-RCO IQ3_S GGUF at 3.5 bits/weight. Authors report 99.8% of BF16 on task average. Details and caveats under Quality.


Models

Same flags, same budget for both. Pick one:

buildrepofilesize
StockISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUFQwen3.8-27B-GSQ-RCO-IQ3_S.gguf11.0 GiB
Abliterated0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUFRVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S.gguf11.0 GiB
Abliterated + MTP head (what this box runs)same repoRVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf11.40 GiB

Measured on the same card, same protocol: 3 requests of 10801 tokens, 256-token replies. These are 11k prompts, so VRAM peaks below the 57k figure above:

metricStock IQ3_SAbliterated -mtp
Prefill600 tok/s564 tok/s (-6%)
Decode16.8 tok/s15.6 tok/s (-7%)
VRAM idle to peak11724 to 11932 MiB11758 to 11758 MiB
Free at peak22 MiB196 MiB
RSS peak3.9 GiB4.3 GiB

The abliterated file is the ISTA quant with weights edited in place. All 851 tensors match stock except 48 ssm_alpha tensors at F32 instead of BF16 (+22.5 MiB), plus one extra block blk.64 (draft head, 15 tensors, 430 MiB). At -ngl 58 that leaves 7 blocks on the host instead of 6. That accounts for most of the 7% slowdown: moving block 58 to CPU on stock (-ngl 57) costs the same, 16.7 vs 17.7 tok/s. Idle cost is +34 MiB, so the budget below covers both.


Hardware

Reference box: Ryzen 9 9900X 12C/24T, RTX 4070 SUPER 12 GB, 30 GB RAM (8 GB is enough), NVMe, Linux, GCC 16.2.1, CUDA 13.4.

Throughput follows GPU memory bandwidth (plus CPU bandwidth for the host-resident layers). Placement comes from the architecture and quant, so it carries over to other 12 GB cards.


Why it fits

1. Only 16 of 64 layers need KV cache

Qwen3.8-27B is hybrid SSM + attention (general.architecture = qwen35):

layer typecountKV cache
SSM / linear attention48no, recurrent state only
Full attention (3, 7, 11, ... 63)16yes, every 4th layer

An all-attention model with these dims would need about 16 GB of f16 KV at 64k. This one needs 1.1 GB:

16 layers x 4 KV heads x 256 head_dim x 2 (K+V) = 32768 elements/token
q4_0 = 18 bytes / 32 elements                   = 18432 bytes/token
x 64000 tokens                                  = 1125 MiB
KV typeat 64kfits
f164000 MiBno
q8_02125 MiBno
q4_01125 MiByes

-ctk q4_0 -ctv q4_0 is what makes 64k load.

2. 11 GiB of weights, split over PCIe

IQ3_S averages 3.5 bits/weight. File is 11771546784 bytes (11.0 GiB). Sizes from the GGUF header:

tensor groupsizewhere
Blocks 0-579072 MiBGPU (-ngl 58)
output.weight + output_norm682 MiBGPU
Blocks 58-631073 MiBCPU
token_embd.weight388 MiBCPU (-ot token_embd=CPU)

Embedding lookup runs once per request and stays mmapped, so offloading it saves 388 MiB for no measurable cost.

3. Total budget

componentMiB
Weights on GPU (blocks 0-57 + output)9754
KV cache (q4_0, full 64k)1125
CUDA context + graph (idle residual)~845
Measured idle11724
Measured peak (first-use workspace + prompt cache)11947
Free at peak7

Two notes on that last row. 12282 - used overstates free; the driver holds about 330 MiB outside it. Peak is what matters. Observed peaks were 11916-11947 MiB (7-38 MiB free) depending on request. It never OOMed here, but there is no room left for larger -ub, a second slot, or longer context.

  • -ngl 64 OOMs. Blocks 58-63 add 1073 MiB: 10827 + 1125 KV = 11952 MiB before context, graphs, and driver overhead, about 500 MiB over. It fails on KV alloc: allocating 1125.00 MiB on device 0: cudaMalloc failed: out of memory then failed to allocate buffer for kv cache.
  • -ot token_embd=CPU is required to start at all.

RAM

CPU blocks plus embedding table are 1461 MiB of weights, mmapped. RSS: 2.1-2.2 GiB idle, 2.8 GiB after one 57k request, 4.0 GiB after three 10801-token requests (3.9-4.3 GiB across builds). Grows about 0.6 GiB per cached prompt. That is prompt cache, not weights. Budget 8 GB for a server holding several long contexts. At idle smaps_rollup shows 1598 MiB file-backed plus 525 MiB anonymous.

Load shows a brief VmRSS spike to about 11.4 GiB for a second while the GGUF is mapped, then it drops. At that point Pss_File is 11.1 GiB vs Pss_Anon 0.24 GiB, and system MemAvailable falls only about 2.1 GiB. These are clean page-cache pages. 8 GB RAM loads it.


Quality

BF16 is 53.8 GB. On 12 GB VRAM you quantize regardless. Question is which quant.

GSQ-RCO (GSQ, RCO, ISTA DASLab) is non-uniform: each tensor gets its own quant type from a gradient search under a size budget. Sensitive tensors keep precision, redundant ones lose it.

Published numbers, IQ3_S vs BF16 (sizes decimal GB; the 11.8 GB row is the 11.0 GiB file used here):

variantbpwsizeWiki pplZS avgRecoveryAIME25GPQA-DLCB v6
BF1616.0053.8 GB7.0574.34100.0%100.0089.9085.71
GSQ-RCO IQ3_S (this recipe)3.5011.8 GB7.0774.47100.2%100.0089.3985.71
UD-IQ3_S (uniform ref)3.5212.0 GB7.1675.49101.5%96.6789.9084.00

Task average 91.70 vs 91.87. Ties AIME25 and LiveCodeBench, -0.51 on GPQA-Diamond.

Caveats:

  • These are the release authors' benchmarks. Little independent verification.
  • Still 3.5-bit, not original. Perplexity is worse: wiki 7.07 vs 7.05 (+0.3%), C4 11.76 vs 11.45 (+2.7%), FineWeb-Edu 8.34 vs 8.14 (+2.5%). Exact-match ties can hide loss in recall, style, and long-context behavior.
  • Best option at this size anyway. For about 11 GiB you get near-BF16 (GSQ-RCO) or a clearly worse uniform quant. No BF16-shaped option exists on 12 GB.

Reproduce

1. Model

# Stock, IQ3_S (11.0 GiB)
hf download ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF \
  Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf --local-dir ~/models/qwen3.8-27b

# Abliterated, same quant (11.0 GiB; -mtp twin 11.40 GiB)
hf download 0bserverx/Qwen3.8-27B-Heretic-GSQ-RCO-GGUF \
  RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S.gguf --local-dir ~/models/qwen3.8-27b

Name the file exactly. --include "*IQ3_S*" also pulls the -mtp twin. Other sizes in the ISTA repo: IQ3_XXS (10.1 GB, smaller/faster), IQ2_XS / IQ2_S (8.4 and 9.3 GB, clear quality cost). mmproj adds vision; this recipe is text-only.

Both builds use the same flags below. Any other qwen35 Qwen3.8-27B GGUF up to about 11 GiB is a drop-in.

2. Build

./build.sh        # pins commit 254b177, flags in the script

That commit is clean upstream. Pin matters: buffer sizes and --fit change between releases, so retune -ngl if you move.

3. Serve

./serve.sh                                       # defaults, or:
MODEL=/path/to/model.gguf PORT=8105 ./serve.sh

Verbatim:

llama-server \
  -m  ~/models/qwen3.8-27b/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf \
  --host 127.0.0.1 --port 8105 \
  --parallel 1 -c 64000 -b 2048 -ub 96 -t 12 \
  -ctk q4_0 -ctv q4_0 \
  -ot token_embd=CPU \
  -ngl 58

Abliterated differs only in -m:

  -m  ~/models/qwen3.8-27b/RVN-Qwen3.8-27B-Heretic-GSQ-RCO-IQ3_S-mtp.gguf
flagreason
-ngl 5858 of 64 blocks on GPU = 9072 MiB. 64 OOMs. Drop by one on smaller cards
-ot token_embd=CPU388 MiB embedding table off GPU
-ctk q4_0 -ctv q4_0KV 1125 MiB instead of 4000. Needed for 64k. Third-party A/Bs on other models show about +0.4% perplexity; not remeasured here
-c 64000Sized to that KV budget
--parallel 1One slot. KV allocated once; extra clients queue instead of using more memory
-b 2048 -ub 96Physical batch 96 keeps the graph inside the ~845 MiB residual. Higher is faster prefill, less headroom
-t 12Threads for host layers. Half the physical cores; SMT siblings measured worse
--host 127.0.0.1Loopback only. Add --api-key and set CORS if you expose it

Flash attention stays at default auto, which is on with quantized V cache. -fa off -ctv q4_0 fails with quantized V cache requires flash_attn to be enabled. No -fa flag is needed. CUDA graphs are compiled in (GGML_CUDA_GRAPHS=ON, upstream default OFF). The server logs graphs reused per request when active.


What did not fit

  • MTP speculative decoding (--spec-type draft-mtp). Fails two ways. At -c 64000 it never starts: failed to allocate CUDA0 buffer of size 387146240. At -c 32000 it loads then OOMs on the first request (peak 11862 MiB, 92 free). The draft wants its own context, KV, and graph buffers on top of a full allocation. Carrying the -mtp head without --spec-type is fine (+34 MiB).
  • f16 KV. 4000 MiB at 64k. Does not fit.
  • q8_0 KV. 2125 MiB at 64k, i.e. +1000 MiB over q4_0 against single-digit free. Only works around -c 32000 (KV 1063 MiB) or lower -ngl.

Measuring it yourself

  • Decode/prefill: print_timing lines in the server log, e.g. eval time = 47637.25 ms / 809 tokens (58.96 ms per token, 16.96 tokens per second). Decode depends on context: 18.1-18.4 tok/s on a 30-token prompt, 17.0 on 4-11k, 16.8 across 163 requests, 13.5 at 57k.
  • VRAM: sample peak during a request, not idle: nvidia-smi --query-gpu=memory.used,memory.free --format=csv once per second.
  • Tensor sizes: read the GGUF header. 851 tensors with type and dims. Bytes = elements x bits/8 / block_size; consecutive data offsets give the same answer. Keys: qwen35.block_count, qwen35.attention.head_count_kv, qwen35.attention.key_length, qwen35.full_attention_interval.

License

MIT for these scripts, see LICENSE. Models are Apache-2.0 per Hugging Face metadata.