[!IMPORTANT] This is the PrismML fork of llama.cpp, the main line behind the Bonsai models (branch
prism, developed asprism-v7). It tracks current mainline llama.cpp and adds the fork's low-bit formats and runtime features on top.New here? Start with the Bonsai-demo repo. It downloads the right models and the correct prebuilt binaries for your hardware/backend automatically.
Which ternary model file to use:
*-PQ2_0.gguf(fork group-128, ggml id 142): preferred on Metal, CUDA, HIP and CPU. About 6% smaller than group-64.*-Q2_0_g64.gguf/ 27B*-Q2_g64.gguf(official group-64, ggml id 42): runs on every backend here AND on mainline llama.cpp. If unsure, use this. Newer model releases name this file plain*-Q2_0.gguf.*-Q2_0.ggufon OLDER model repos is the deprecated legacy format (group 128 stored as id 42). It does not load on these builds; the error tells you which file to get instead. If you must run it, use the frozenprism-v5line and its final releaseprism-b9601.Speculative decoding (dspark) is supported via mainline's draft-dspark plus fork patches. Drafters published for older model releases need a one-time conversion with
gguf-dspark-to-dflash(see SPECULATIVE.md in Bonsai-demo); newer releases ship ready-to-use drafters.Do NOT build from
prism-v6(stale mid-migration snapshot) and do NOT mix this fork'sggml-*libraries with a stock llama.cpp build.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT license(top 30 of 443)
C++
55.7%
C
15.5%
Python
7.2%
Cuda
5.7%
TypeScript
4.3%
Svelte
2.3%
HTML
2.1%
Metal
1.6%
Jinja
1.2%
[!IMPORTANT] This is the PrismML fork of llama.cpp, the main line behind the Bonsai models (branch
prism, developed asprism-v7). It tracks current mainline llama.cpp and adds the fork's low-bit formats and runtime features on top.New here? Start with the Bonsai-demo repo. It downloads the right models and the correct prebuilt binaries for your hardware/backend automatically.
Which ternary model file to use:
*-PQ2_0.gguf(fork group-128, ggml id 142): preferred on Metal, CUDA, HIP and CPU. About 6% smaller than group-64.*-Q2_0_g64.gguf/ 27B*-Q2_g64.gguf(official group-64, ggml id 42): runs on every backend here AND on mainline llama.cpp. If unsure, use this. Newer model releases name this file plain*-Q2_0.gguf.*-Q2_0.ggufon OLDER model repos is the deprecated legacy format (group 128 stored as id 42). It does not load on these builds; the error tells you which file to get instead. If you must run it, use the frozenprism-v5line and its final releaseprism-b9601.Speculative decoding (dspark) is supported via mainline's draft-dspark plus fork patches. Drafters published for older model releases need a one-time conversion with
gguf-dspark-to-dflash(see SPECULATIVE.md in Bonsai-demo); newer releases ship ready-to-use drafters.Do NOT build from
prism-v6(stale mid-migration snapshot) and do NOT mix this fork'sggml-*libraries with a stock llama.cpp build.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT license(top 30 of 443)
C++
55.7%
C
15.5%
Python
7.2%
Cuda
5.7%
TypeScript
4.3%
Svelte
2.3%
HTML
2.1%
Metal
1.6%
Jinja
1.2%