jeva.cpp - a llama.cpp fork with a JEV-compatible decision API for all LLMs supported by llama.cpp.
C++
0
9,345 commits
updated Sep 28, 2026
jeva.cpp is a fork of llama.cpp that adds a JEV-compatible decision API to llama-server, enabling Choice, Score and Noul evaluations directly from model logits while preserving standard autoregressive generation.
jeva.cpp is designed to work with all models and platforms supported by llama.cpp, reusing its existing model implementations and inference backends. The JEV decision API requires models that provide next-token vocabulary logits; other model types retain their original functionality.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT licenseC++
55.9%
C
16.2%
Python
7.3%
Cuda
5.4%
TypeScript
4.1%
Svelte
2.1%
HTML
2.0%
Metal
1.6%
Jinja
1.2%
jeva.cpp - a llama.cpp fork with a JEV-compatible decision API for all LLMs supported by llama.cpp.
C++
0
9,345 commits
updated Sep 28, 2026
jeva.cpp is a fork of llama.cpp that adds a JEV-compatible decision API to llama-server, enabling Choice, Score and Noul evaluations directly from model logits while preserving standard autoregressive generation.
jeva.cpp is designed to work with all models and platforms supported by llama.cpp, reusing its existing model implementations and inference backends. The JEV decision API requires models that provide next-token vocabulary logits; other model types retain their original functionality.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
llama.cpp repo and merge PRs into the master branchllama-server - MIT licenseC++
55.9%
C
16.2%
Python
7.3%
Cuda
5.4%
TypeScript
4.1%
Svelte
2.1%
HTML
2.0%
Metal
1.6%
Jinja
1.2%