lkarlslund/laya.cpp

C++ inference for Laya typed decisions - supports CUDA, Vulkan, Core ML, CPU

C++

113

252 commits

updated Sep 27, 2026

See the code

See what people are saying

SourceMessageScoreDate

LibLayaX: run the Laya AI decision model inside your own app (r/LLMDevs)

Dear LLM developers community, I have just released four open-source projects today that let an application use the Laya model directly, with no server and no Python. **What Laya is?** If you build agents, this is the step where you ask a big model "should I route this to billing?" and wait a…

6

Oct 3, 2026

README

laya.cpp

Native C++ inference for Laya typed decisions, powered by ggml with CUDA and Vulkan and Apple Core ML backends. Tokenization, inference and JSON output run without Python. Supports the english, multilingual and typed-decisions models, plus a JEV-compatible HTTP server with automatic request batching.

Binary releases provide Windows and Linux x64 CUDA/Vulkan executables plus a macOS arm64 Core ML executable. See runtime requirements to choose CUDA 12 or 13, check drivers, Windows CUDA DLLs and compiled Core ML model buckets.

Performance

Latest paired comparisons against matching-precision Python, using 250 fixed questions across all three models at batch sizes 1, 2, 4 and 8. Higher is better; 1× means equal throughput. NVIDIA measurements use an RTX PRO 6000 Blackwell (96 GB), capped at 450 W; AMD measurements use a Radeon 8060S.

GPUBackend / modeThroughput relative to Python
NVIDIACUDA optimized FP321.35–2.48×
NVIDIACUDA BF161.11–2.71×
NVIDIAVulkan plain FP320.44–0.73×
NVIDIAVulkan compensated FP320.60–1.14×
NVIDIAVulkan FP16 / BF160.64–1.02×
AMDVulkan plain FP320.61–1.82×
AMDVulkan compensated FP321.26–2.45×
AMDVulkan FP160.72–1.25×
AMDVulkan BF160.49–0.92×

All measured answer checks pass: exact categories and numeric absolute error at most 0.0001. Timings include preprocessing, inference and formatting, excluding model loading and JSON transport.

See latest measurements for per-model throughput, Vulkan FP32 results, direct CUDA/Vulkan comparisons and measurement identities.

Build

Requires a C++20 compiler, CMake 3.24+, ICU and nlohmann-json, plus the chosen GPU backend's dependencies. On Debian-like systems, install libicu-dev and nlohmann-json3-dev for the host dependencies.

git submodule update --init --recursive
cmake -S . -B build-cuda -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CUDA_ARCHITECTURES=120
cmake --build build-cuda --parallel 8

Architecture 120 targets RTX Blackwell; select the architecture for your GPU. For Vulkan, install the Vulkan loader/headers, glslc and SPIR-V headers, then:

cmake -S . -B build-vulkan -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DLAYA_CUDA=OFF -DLAYA_VULKAN=ON
cmake --build build-vulkan --parallel 8

On Apple Silicon, Core ML uses the CPU, GPU and Neural Engine through MLComputeUnitsAll. The build requires macOS 12 or newer, full Xcode (the Command Line Tools alone are not sufficient), and the host dependencies:

brew install cmake ninja icu4c nlohmann-json
scripts/build_coreml.sh

The script selects the standard /Applications/Xcode.app, locates the Homebrew packages, initializes submodules, configures Release for arm64 and macOS 12, and builds build-coreml/bin/laya-cli. Pass --test to run CTest, or --fresh to discard a stale CMake cache. Run --help for all options.

The checkpoint must first be exported and compiled. See Core ML for model preparation, manual build commands, bucket choices and M1 validation.

Run

Download the models with the optional Python tooling, then run native inference:

python scripts/download_model.py --variant all
build-cuda/bin/laya-cli --model models/laya --variant english \
  --tensor-core-fp32 --flash-fp32 --input benchmarks/cases/smoke.json

Choose --variant multilingual or --variant typed-decisions for another model. Requests that exceed the selected checkpoint's token budgets are rejected by default. Pass --allow-truncation to retain the older behavior for clients that depend on silently shortened input. For Vulkan, use build-vulkan/bin/laya-cli --vulkan and omit --flash-fp32. Strict FP32 is the default; --tensor-core-fp32 enables compensated FP32 projections. Both GPU backends support --bf16; Vulkan also supports --fp16. See precision and Vulkan support for tested hardware and build requirements.

For Core ML, run build-coreml/bin/laya-cli --coreml; model precision is fixed at export time, so do not combine it with --fp16 or --bf16.

Without --input, the CLI accepts one JSON request or request array per line:

{"state":"Please refund the duplicate charge.","questions":{"refund":{"type":"noul","instructions":"Does the customer ask for a refund?"}}}

To serve HTTP on 127.0.0.1:8080:

build-cuda/bin/laya-cli --server --port 8080 --variant english \
  --tensor-core-fp32 --flash-fp32

The JEV-compatible endpoint is POST /v1/systemone. Concurrent requests are batched automatically. See HTTP serving for examples and settings.

Documentation

License

MIT. Dependencies and model files retain their own licenses.

Significant stargazers

r33drichards

34 followers · starred Sep 2026

Alexandre Espinosa Menor

73 followers · starred Sep 2026

Maksim Soltan

211 followers · starred Sep 2026

João Paulo Silva de Souza

41 followers · starred Sep 2026

lkarlslund/laya.cpp

C++ inference for Laya typed decisions - supports CUDA, Vulkan, Core ML, CPU

C++

113

252 commits

updated Sep 27, 2026

See the code

See what people are saying

SourceMessageScoreDate

LibLayaX: run the Laya AI decision model inside your own app (r/LLMDevs)

Dear LLM developers community, I have just released four open-source projects today that let an application use the Laya model directly, with no server and no Python. **What Laya is?** If you build agents, this is the step where you ask a big model "should I route this to billing?" and wait a…

6

Oct 3, 2026

README

laya.cpp

Native C++ inference for Laya typed decisions, powered by ggml with CUDA and Vulkan and Apple Core ML backends. Tokenization, inference and JSON output run without Python. Supports the english, multilingual and typed-decisions models, plus a JEV-compatible HTTP server with automatic request batching.

Binary releases provide Windows and Linux x64 CUDA/Vulkan executables plus a macOS arm64 Core ML executable. See runtime requirements to choose CUDA 12 or 13, check drivers, Windows CUDA DLLs and compiled Core ML model buckets.

Performance

Latest paired comparisons against matching-precision Python, using 250 fixed questions across all three models at batch sizes 1, 2, 4 and 8. Higher is better; 1× means equal throughput. NVIDIA measurements use an RTX PRO 6000 Blackwell (96 GB), capped at 450 W; AMD measurements use a Radeon 8060S.

GPUBackend / modeThroughput relative to Python
NVIDIACUDA optimized FP321.35–2.48×
NVIDIACUDA BF161.11–2.71×
NVIDIAVulkan plain FP320.44–0.73×
NVIDIAVulkan compensated FP320.60–1.14×
NVIDIAVulkan FP16 / BF160.64–1.02×
AMDVulkan plain FP320.61–1.82×
AMDVulkan compensated FP321.26–2.45×
AMDVulkan FP160.72–1.25×
AMDVulkan BF160.49–0.92×

All measured answer checks pass: exact categories and numeric absolute error at most 0.0001. Timings include preprocessing, inference and formatting, excluding model loading and JSON transport.

See latest measurements for per-model throughput, Vulkan FP32 results, direct CUDA/Vulkan comparisons and measurement identities.

Build

Requires a C++20 compiler, CMake 3.24+, ICU and nlohmann-json, plus the chosen GPU backend's dependencies. On Debian-like systems, install libicu-dev and nlohmann-json3-dev for the host dependencies.

git submodule update --init --recursive
cmake -S . -B build-cuda -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CUDA_ARCHITECTURES=120
cmake --build build-cuda --parallel 8

Architecture 120 targets RTX Blackwell; select the architecture for your GPU. For Vulkan, install the Vulkan loader/headers, glslc and SPIR-V headers, then:

cmake -S . -B build-vulkan -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DLAYA_CUDA=OFF -DLAYA_VULKAN=ON
cmake --build build-vulkan --parallel 8

On Apple Silicon, Core ML uses the CPU, GPU and Neural Engine through MLComputeUnitsAll. The build requires macOS 12 or newer, full Xcode (the Command Line Tools alone are not sufficient), and the host dependencies:

brew install cmake ninja icu4c nlohmann-json
scripts/build_coreml.sh

The script selects the standard /Applications/Xcode.app, locates the Homebrew packages, initializes submodules, configures Release for arm64 and macOS 12, and builds build-coreml/bin/laya-cli. Pass --test to run CTest, or --fresh to discard a stale CMake cache. Run --help for all options.

The checkpoint must first be exported and compiled. See Core ML for model preparation, manual build commands, bucket choices and M1 validation.

Run

Download the models with the optional Python tooling, then run native inference:

python scripts/download_model.py --variant all
build-cuda/bin/laya-cli --model models/laya --variant english \
  --tensor-core-fp32 --flash-fp32 --input benchmarks/cases/smoke.json

Choose --variant multilingual or --variant typed-decisions for another model. Requests that exceed the selected checkpoint's token budgets are rejected by default. Pass --allow-truncation to retain the older behavior for clients that depend on silently shortened input. For Vulkan, use build-vulkan/bin/laya-cli --vulkan and omit --flash-fp32. Strict FP32 is the default; --tensor-core-fp32 enables compensated FP32 projections. Both GPU backends support --bf16; Vulkan also supports --fp16. See precision and Vulkan support for tested hardware and build requirements.

For Core ML, run build-coreml/bin/laya-cli --coreml; model precision is fixed at export time, so do not combine it with --fp16 or --bf16.

Without --input, the CLI accepts one JSON request or request array per line:

{"state":"Please refund the duplicate charge.","questions":{"refund":{"type":"noul","instructions":"Does the customer ask for a refund?"}}}

To serve HTTP on 127.0.0.1:8080:

build-cuda/bin/laya-cli --server --port 8080 --variant english \
  --tensor-core-fp32 --flash-fp32

The JEV-compatible endpoint is POST /v1/systemone. Concurrent requests are batched automatically. See HTTP serving for examples and settings.

Documentation

License

MIT. Dependencies and model files retain their own licenses.

Significant stargazers

r33drichards

34 followers · starred Sep 2026

Alexandre Espinosa Menor

73 followers · starred Sep 2026

Maksim Soltan

211 followers · starred Sep 2026

João Paulo Silva de Souza

41 followers · starred Sep 2026

Languages

C++

81.6%

Python

10.6%

CMake

4.3%

Cuda

2.0%