On Debian/Ubuntu ARM, install GCC/G++ if needed:
sudo apt update
sudo apt install -y build-essential
git clone git@github.com:deepgrove-ai/llama.cpp.git
cd llama.cpp
rm -rf ./build
uvx cmake -B build \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_METAL=OFF
uvx cmake --build build -j
macOS uses Apple Clang. uvx provides CMake, so no separate CMake installation is needed.
uvx --from huggingface_hub hf download \
deepgrove/maple-preview-GGUF \
maple-preview-TQ2_0-head-Q4_K.gguf \
--local-dir models
Hugging Face: deepgrove/maple-preview-GGUF
| Variant | GGUF size |
|---|---|
| TQ1_0 + Q4_K head | 4.64 GiB |
| TQ1_0 + FP16 head | 5.06 GiB |
| TQ2_0 + Q4_K head | 5.50 GiB |
| TQ2_0 + FP16 head | 5.91 GiB |
TQ1_0 and TQ2_0 are different ternary packing schemes. Use TQ2_0 for generally faster speeds but slightly higher memory footprint.
./build/bin/llama-completion \
-m models/maple-preview-TQ2_0-head-Q4_K.gguf \
--threads 16 \
--temp 1.0 \
--top-p 0.95 \
--jinja \
--conversation
--jinja applies Maple's embedded chat template exactly, including its thinking prefix.
M5 Pro, CPU-only, 16 threads, 512 prompt tokens, 128 generated tokens, 3 repetitions:
| Matrix weights | LM head | GGUF size | Prefill (pp512) | Decode (tg128) |
|---|---|---|---|---|
| TQ1_0 | FP16 | 5.06 GiB | 515.41 ± 0.28 tokens/s | 161.06 ± 0.57 tokens/s |
| TQ1_0 | Q4_K | 4.64 GiB | 513.33 ± 3.62 tokens/s | 231.13 ± 0.13 tokens/s |
| TQ2_0 | FP16 | 5.91 GiB | 618.57 ± 2.12 tokens/s | 169.81 ± 2.94 tokens/s |
| TQ2_0 | Q4_K | 5.50 GiB | 610.48 ± 3.76 tokens/s | 252.74 ± 0.37 tokens/s |
Adjust -t for #of threads.
FP16 LM head:
./build/bin/llama-bench \
-m models/maple-preview-TQ2_0-head-F16.gguf \
-p 512 \
-n 128 \
-r 3 \
-t 16 \
-dev none \
-ngl 0
Q4_K LM head:
./build/bin/llama-bench \
-m models/maple-preview-TQ2_0-head-Q4_K.gguf \
-p 512 \
-n 128 \
-r 3 \
-t 16 \
-dev none \
-ngl 0
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / dev branches / compile times / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT license(top 30 of 447)
C++
55.4%
C
16.3%
Python
7.1%
Cuda
5.6%
TypeScript
4.0%
HTML
2.3%
Svelte
2.2%
Metal
1.4%
Jinja
1.2%
GLSL
1.0%
On Debian/Ubuntu ARM, install GCC/G++ if needed:
sudo apt update
sudo apt install -y build-essential
git clone git@github.com:deepgrove-ai/llama.cpp.git
cd llama.cpp
rm -rf ./build
uvx cmake -B build \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_METAL=OFF
uvx cmake --build build -j
macOS uses Apple Clang. uvx provides CMake, so no separate CMake installation is needed.
uvx --from huggingface_hub hf download \
deepgrove/maple-preview-GGUF \
maple-preview-TQ2_0-head-Q4_K.gguf \
--local-dir models
Hugging Face: deepgrove/maple-preview-GGUF
| Variant | GGUF size |
|---|---|
| TQ1_0 + Q4_K head | 4.64 GiB |
| TQ1_0 + FP16 head | 5.06 GiB |
| TQ2_0 + Q4_K head | 5.50 GiB |
| TQ2_0 + FP16 head | 5.91 GiB |
TQ1_0 and TQ2_0 are different ternary packing schemes. Use TQ2_0 for generally faster speeds but slightly higher memory footprint.
./build/bin/llama-completion \
-m models/maple-preview-TQ2_0-head-Q4_K.gguf \
--threads 16 \
--temp 1.0 \
--top-p 0.95 \
--jinja \
--conversation
--jinja applies Maple's embedded chat template exactly, including its thinking prefix.
M5 Pro, CPU-only, 16 threads, 512 prompt tokens, 128 generated tokens, 3 repetitions:
| Matrix weights | LM head | GGUF size | Prefill (pp512) | Decode (tg128) |
|---|---|---|---|---|
| TQ1_0 | FP16 | 5.06 GiB | 515.41 ± 0.28 tokens/s | 161.06 ± 0.57 tokens/s |
| TQ1_0 | Q4_K | 4.64 GiB | 513.33 ± 3.62 tokens/s | 231.13 ± 0.13 tokens/s |
| TQ2_0 | FP16 | 5.91 GiB | 618.57 ± 2.12 tokens/s | 169.81 ± 2.94 tokens/s |
| TQ2_0 | Q4_K | 5.50 GiB | 610.48 ± 3.76 tokens/s | 252.74 ± 0.37 tokens/s |
Adjust -t for #of threads.
FP16 LM head:
./build/bin/llama-bench \
-m models/maple-preview-TQ2_0-head-F16.gguf \
-p 512 \
-n 128 \
-r 3 \
-t 16 \
-dev none \
-ngl 0
Q4_K LM head:
./build/bin/llama-bench \
-m models/maple-preview-TQ2_0-head-Q4_K.gguf \
-p 512 \
-n 128 \
-r 3 \
-t 16 \
-dev none \
-ngl 0
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / dev branches / compile times / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT license(top 30 of 447)
C++
55.4%
C
16.3%
Python
7.1%
Cuda
5.6%
TypeScript
4.0%
HTML
2.3%
Svelte
2.2%
Metal
1.4%
Jinja
1.2%
GLSL
1.0%