LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Two opt-in optimizations for large MoE models whose experts are offloaded to system RAM
(--n-cpu-moe), found and implemented by Fable. Both are off by default, toggled via
environment variables, and produce token-identical output to mainline.
| Env var | What it does |
|---|---|
GGML_CUDA_REGISTER_HOST=1 | Page-locks (pins) the mmap'd CPU expert weights so host→device copies go straight over DMA instead of through the driver's hidden bounce buffer (~6–7 → ~20 GB/s). |
GGML_SCHED_PREFETCH_EXPERTS=1 | Prefetches each layer's experts on a second CUDA stream, so the weight uploads overlap compute instead of stalling the GPU. |
Measured on an RTX 3060 12GB with Qwen3.6-35B-A3B (--n-cpu-moe 26), prompt-processing at 2048 (MODEL = path to your .gguf):
# baseline (patches off):
./build/bin/llama-bench -m MODEL -ngl 99 -ncmoe 26 -p 2048 -n 0 -r 5 -b 2048 -ub 2048
# patched (both optimizations on):
GGML_CUDA_REGISTER_HOST=1 GGML_SCHED_PREFETCH_EXPERTS=1 \
./build/bin/llama-bench -m MODEL -ngl 99 -ncmoe 26 -p 2048 -n 0 -r 5 -b 2048 -ub 2048
Result: ~1143 → ~1880 t/s prefill (+64%) — same GPU, same settings, token-identical.
Branches: fable5/host-register (pinning only) · fable5/prefetch-experts (both — this branch).
For MoE models whose routed experts live in system RAM (--n-cpu-moe), this fork can keep the
most-frequently-routed experts of each layer resident in VRAM. Decode runs the hot experts on
GPU and only the cold remainder on CPU; the two halves are merged exactly, so output is
bit-identical to baseline. Opt-in, off by default.
Measured on an RTX 3060 12GB (-ngl 99 -ncmoe 99 -fa 1):
| Model | Baseline tg | Cached tg | Prefill |
|---|---|---|---|
| Qwen3.6-35B-A3B Q4_K_M (256 experts/layer) | 42.3 | 51.3 (+21%) @ 124 slots | +14% |
— same, stacked with --spec-type draft-mtp | 41.7 | 69.3 (+66%) @ 112 slots | — |
| — same, plus async CPU splits (default on) | 41.7 | 74.2 (+78%) @ 88 slots | — |
| GLM-4.7-Flash Q4_K_M (64 experts/layer) | 32.1 | 46.3 (+44%) @ 40 slots | +64% |
| Laguna-S-2.1-118B-A8B IQ4_XS (256 experts/layer) | 11.5 | 12.1 (+5%) @ 36 slots | +12% |
| Qwen3.8-Flash-Next 177B UD-IQ3_XXS (512 experts/layer, 10 active) | 16.6 | 24.4 (+47%) @ 56 slots, with --spec-type draft-mtp | — |
Supported architectures: qwen35moe, qwen4exp (Qwen3.8-Flash-Next), deepseek2, laguna (plain fused-SILU gated expert FFN,
separate gate/up/down tensors). Other architectures run unchanged.
1. Capture a routing profile (one time per model — records which experts the router picks):
MOE_TRACE_OUT=mymodel-code.csv ./build/bin/llama-moe-trace -m model.gguf \
-ngl 99 -ncmoe 99 -fa 1 -c 4096 -n 512 -p "<a code-flavored prompt>"
MOE_TRACE_OUT=mymodel-chat.csv ./build/bin/llama-moe-trace -m model.gguf \
-ngl 99 -ncmoe 99 -fa 1 -c 4096 -n 512 -p "<a chat-flavored prompt>"
cat mymodel-code.csv mymodel-chat.csv > mymodel-merged.csv
512 generated tokens per prompt is enough. Merge traces from contrasting workloads — a merged profile measures within 1% of per-workload specialist profiles, so one merged CSV per model is all you need.
2. Serve with the cache:
./build/bin/llama-server -m model.gguf -ngl 99 -ncmoe 99 -fa 1 \
--moe-cache-profile mymodel-merged.csv --moe-cache-slots 112
Also works per model in a --models-preset INI section (moe-cache-profile = ...,
moe-cache-slots = ...), and as env vars LLAMA_ARG_MOE_CACHE_PROFILE / LLAMA_ARG_MOE_CACHE_SLOTS
(or legacy GGML_MOE_CACHE_PROFILE / GGML_MOE_CACHE_SLOTS, which llama-bench also accepts).
3. Confirm it engaged — look for this line at load:
init_moe_expert_cache: expert cache: 40 layers x 112 slots, 8164.00 MiB uploaded to CUDA0
A warning instead of this line means the cache fell back to baseline (see Tuning).
--moe-cache-slots is the main knob — experts cached per layer. Throughput rises with slot
count until the pack no longer fits in VRAM. The pack is all-or-nothing: an oversized request
logs pack allocation failed - expert cache disabled and runs at baseline speed (it does not
partially fill). The warning reports the per-slot cost and the maximum count that could fit —
set slots to that, minus headroom for KV/compute buffers which allocate afterwards.--no-sched-async-cpu to disable;
llama-bench --sched-async-cpu 0,1 benches both). Worth +4-5% with speculative decoding, ~±2%
without it; outputs stay bit-identical either way.--n-cpu-moe).
Pure -ncmoe 99 + max slots is the simple default; a hybrid (e.g. -ncmoe 30 + fewer slots)
buys ~1% decode and ~3% prefill at best.-c and shrinks the viable slot count.
Compressing the KV cache (-ctk/-ctv, e.g. TurboQuant types) frees VRAM that converts
directly into slots — often worth more than the KV precision costs.--spec-type draft-mtp composes with the
cache (+48% cache × +12% MTP ≈ +66% on Qwen); reserve ~1 GB for the draft context by dropping
a few slots.| Symptom | Cause |
|---|---|
cannot open profile '...' | Path not visible to the process (e.g. not mounted into the container). |
pack allocation failed | Slot count too big — read the fit math in the warning and reduce. |
no CPU-resident MoE layers | Experts are already on GPU (no --n-cpu-moe) — nothing to cache. |
| No init line, no warning | Architecture not wired for the cache — model runs unchanged. |
| Model loads, then context creation OOMs | Pack fits but KV/compute don't — drop a few slots or shrink/compress KV. |
-hf are now stored in the standard Hugging Face cache directory, enabling sharing with other HF tools.gpt-oss model with native MXFP4 format has been added | PR | Collaboration with NVIDIA | Commentllama-server: #12898 | documentationA few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT licenseC++
56.9%
C
14.3%
Python
7.5%
Cuda
5.4%
TypeScript
3.6%
HTML
2.5%
Svelte
2.1%
Metal
1.5%
Jinja
1.2%
GLSL
1.1%
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Two opt-in optimizations for large MoE models whose experts are offloaded to system RAM
(--n-cpu-moe), found and implemented by Fable. Both are off by default, toggled via
environment variables, and produce token-identical output to mainline.
| Env var | What it does |
|---|---|
GGML_CUDA_REGISTER_HOST=1 | Page-locks (pins) the mmap'd CPU expert weights so host→device copies go straight over DMA instead of through the driver's hidden bounce buffer (~6–7 → ~20 GB/s). |
GGML_SCHED_PREFETCH_EXPERTS=1 | Prefetches each layer's experts on a second CUDA stream, so the weight uploads overlap compute instead of stalling the GPU. |
Measured on an RTX 3060 12GB with Qwen3.6-35B-A3B (--n-cpu-moe 26), prompt-processing at 2048 (MODEL = path to your .gguf):
# baseline (patches off):
./build/bin/llama-bench -m MODEL -ngl 99 -ncmoe 26 -p 2048 -n 0 -r 5 -b 2048 -ub 2048
# patched (both optimizations on):
GGML_CUDA_REGISTER_HOST=1 GGML_SCHED_PREFETCH_EXPERTS=1 \
./build/bin/llama-bench -m MODEL -ngl 99 -ncmoe 26 -p 2048 -n 0 -r 5 -b 2048 -ub 2048
Result: ~1143 → ~1880 t/s prefill (+64%) — same GPU, same settings, token-identical.
Branches: fable5/host-register (pinning only) · fable5/prefetch-experts (both — this branch).
For MoE models whose routed experts live in system RAM (--n-cpu-moe), this fork can keep the
most-frequently-routed experts of each layer resident in VRAM. Decode runs the hot experts on
GPU and only the cold remainder on CPU; the two halves are merged exactly, so output is
bit-identical to baseline. Opt-in, off by default.
Measured on an RTX 3060 12GB (-ngl 99 -ncmoe 99 -fa 1):
| Model | Baseline tg | Cached tg | Prefill |
|---|---|---|---|
| Qwen3.6-35B-A3B Q4_K_M (256 experts/layer) | 42.3 | 51.3 (+21%) @ 124 slots | +14% |
— same, stacked with --spec-type draft-mtp | 41.7 | 69.3 (+66%) @ 112 slots | — |
| — same, plus async CPU splits (default on) | 41.7 | 74.2 (+78%) @ 88 slots | — |
| GLM-4.7-Flash Q4_K_M (64 experts/layer) | 32.1 | 46.3 (+44%) @ 40 slots | +64% |
| Laguna-S-2.1-118B-A8B IQ4_XS (256 experts/layer) | 11.5 | 12.1 (+5%) @ 36 slots | +12% |
| Qwen3.8-Flash-Next 177B UD-IQ3_XXS (512 experts/layer, 10 active) | 16.6 | 24.4 (+47%) @ 56 slots, with --spec-type draft-mtp | — |
Supported architectures: qwen35moe, qwen4exp (Qwen3.8-Flash-Next), deepseek2, laguna (plain fused-SILU gated expert FFN,
separate gate/up/down tensors). Other architectures run unchanged.
1. Capture a routing profile (one time per model — records which experts the router picks):
MOE_TRACE_OUT=mymodel-code.csv ./build/bin/llama-moe-trace -m model.gguf \
-ngl 99 -ncmoe 99 -fa 1 -c 4096 -n 512 -p "<a code-flavored prompt>"
MOE_TRACE_OUT=mymodel-chat.csv ./build/bin/llama-moe-trace -m model.gguf \
-ngl 99 -ncmoe 99 -fa 1 -c 4096 -n 512 -p "<a chat-flavored prompt>"
cat mymodel-code.csv mymodel-chat.csv > mymodel-merged.csv
512 generated tokens per prompt is enough. Merge traces from contrasting workloads — a merged profile measures within 1% of per-workload specialist profiles, so one merged CSV per model is all you need.
2. Serve with the cache:
./build/bin/llama-server -m model.gguf -ngl 99 -ncmoe 99 -fa 1 \
--moe-cache-profile mymodel-merged.csv --moe-cache-slots 112
Also works per model in a --models-preset INI section (moe-cache-profile = ...,
moe-cache-slots = ...), and as env vars LLAMA_ARG_MOE_CACHE_PROFILE / LLAMA_ARG_MOE_CACHE_SLOTS
(or legacy GGML_MOE_CACHE_PROFILE / GGML_MOE_CACHE_SLOTS, which llama-bench also accepts).
3. Confirm it engaged — look for this line at load:
init_moe_expert_cache: expert cache: 40 layers x 112 slots, 8164.00 MiB uploaded to CUDA0
A warning instead of this line means the cache fell back to baseline (see Tuning).
--moe-cache-slots is the main knob — experts cached per layer. Throughput rises with slot
count until the pack no longer fits in VRAM. The pack is all-or-nothing: an oversized request
logs pack allocation failed - expert cache disabled and runs at baseline speed (it does not
partially fill). The warning reports the per-slot cost and the maximum count that could fit —
set slots to that, minus headroom for KV/compute buffers which allocate afterwards.--no-sched-async-cpu to disable;
llama-bench --sched-async-cpu 0,1 benches both). Worth +4-5% with speculative decoding, ~±2%
without it; outputs stay bit-identical either way.--n-cpu-moe).
Pure -ncmoe 99 + max slots is the simple default; a hybrid (e.g. -ncmoe 30 + fewer slots)
buys ~1% decode and ~3% prefill at best.-c and shrinks the viable slot count.
Compressing the KV cache (-ctk/-ctv, e.g. TurboQuant types) frees VRAM that converts
directly into slots — often worth more than the KV precision costs.--spec-type draft-mtp composes with the
cache (+48% cache × +12% MTP ≈ +66% on Qwen); reserve ~1 GB for the draft context by dropping
a few slots.| Symptom | Cause |
|---|---|
cannot open profile '...' | Path not visible to the process (e.g. not mounted into the container). |
pack allocation failed | Slot count too big — read the fit math in the warning and reduce. |
no CPU-resident MoE layers | Experts are already on GPU (no --n-cpu-moe) — nothing to cache. |
| No init line, no warning | Architecture not wired for the cache — model runs unchanged. |
| Model loads, then context creation OOMs | Pack fits but KV/compute don't — drop a few slots or shrink/compress KV. |
-hf are now stored in the standard Hugging Face cache directory, enabling sharing with other HF tools.gpt-oss model with native MXFP4 format has been added | PR | Collaboration with NVIDIA | Commentllama-server: #12898 | documentationA few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT licenseC++
56.9%
C
14.3%
Python
7.5%
Cuda
5.4%
TypeScript
3.6%
HTML
2.5%
Svelte
2.1%
Metal
1.5%
Jinja
1.2%
GLSL
1.1%