doopeworld/GRIMOIRE

C++

1

341 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

GRIMOIRE: native C++/SYCL LLM inference for Intel Arc Pro B70 — Ornith 1.5 hits 175 tok/s single request, 464 tok/s at concurrency 8 (r/LocalLLM)

https://preview.redd.it/azrcl8q94gth1.png?width=2172&format=png&auto=webp&s=20d4010874d090d1d35ee09a292e77696e4820b0 Hi everyone, I’ve been building **GRIMOIRE**, an experimental LLM inference engine for **Intel Battlemage GPUs**, with the **Arc Pro B70** as its primary target. GRIMOIRE…

2

Oct 4, 2026

README

GRIMOIRE

GRIMOIRE LLM inference banner

LLM inference built for Intel Arc Pro B70 Battlemage GPUs.

GRIMOIRE is a native inference engine written in C++ with SYCL and Level Zero. It is being developed to run and optimize large language models directly on Intel Battlemage hardware, with a focus on the Arc Pro B70.

The aim is straightforward: make high-performance local LLM inference possible on this hardware, using kernels and memory paths designed for the GPU instead of relying on a general-purpose inference stack.

Project status: Experimental. GRIMOIRE is actively developed, but it is not yet a stable, production-ready release. Expect incomplete features, changing model support, and performance or compatibility issues. A ready-to-use inference image or downloadable model package will be added here later.

Why GRIMOIRE was created

Most inference tooling and performance guidance is centered on other GPU platforms. Battlemage owners need an engine that treats Intel hardware as a first-class target and makes its real performance measurable.

GRIMOIRE began as a ground-up inference project for the Arc Pro B70. Building the engine directly exposed important details that a port or a high-level wrapper could hide: how the GPU represents quantized weights, where attention kernels can go wrong, how to keep decode bound by useful memory traffic, and what multi-GPU communication actually costs on this setup.

The project exists to turn those findings into working inference code, reproducible checks, and practical performance improvements for Battlemage users.

What it is for

  • Running supported large language models locally on Intel Arc Pro B70 GPUs.
  • Exploring native C++ / SYCL / Level Zero kernels for model inference.
  • Measuring prefill and token-generation performance on real hardware.
  • Improving single-GPU execution and developing multi-GPU paths for models that need more memory.
  • Checking changes against a growing set of model and checkpoint combinations.

GRIMOIRE is a hardware-focused inference engine, not a hosted model service. Model compatibility and performance depend on the architecture, weight format, configuration, and hardware available.

How it is built

  • C++ inference engine
  • SYCL GPU programming
  • Level Zero device runtime
  • Intel Battlemage, with the Arc Pro B70 as the primary target
  • Custom GPU work for operations such as attention, matrix multiplication, quantized weight handling, and mixture-of-experts layers

The project aims to keep inference native to the target hardware and does not use vLLM, PyTorch, or OpenVINO as its inference backend.

Recent measured results

These are results recorded in the linked project commits on the developer's Arc Pro B70 system. They are examples from specific models and test runs, not universal performance guarantees.

WorkloadRecorded resultContext
Ornith-1.5-35B-A3B token generation198.6 tokens/sSingle B70; 256 generated tokens in the recorded run. Commit
Ornith-1.5-35B-A3B prefill10,030–10,165 tokens/sTwo-GPU pipeline-parallel run, 5,987-token prompt. Output matched the single-B70 run for the checked text. Commit
TP decode communication batching59.3–60.0 tokens/sRecorded TP run after combining independent projection gathers; see the commit for the setup and comparison. Commit

The two-GPU figures demonstrate an actively developed path. Tensor and pipeline parallelism are intended for models that need multiple GPUs for capacity; they are not automatically faster for a model that already fits on one card.

Results change as kernels and test coverage evolve. The commit links include the measurements, validation notes, and limitations for each change. The latest model-zoo audit reports 27 confirmed working model/checkpoint combinations in the local regression sweep, with additional formats and configurations correctly identified as out of scope. Latest audit

Supported models

Everything below is validated end-to-end — loaded, generating coherent text, checked against a reference where one exists — on one Intel Arc Pro B70, by the project's own regression suite. That suite runs after every change that touches shared decode, prefill, attention, or MoE code, not as a one-time check.

ModelFormatsRoleNotes
Ornith-1.5-35B-A3BMXFP4, NVFP4, FP8, GPTQ-Int4, INT4 (AutoRound W4A16), bf16 sourcetarget35B MoE, 256 experts / top-8. 198.6 tok/s decode, 10,030–10,165 tok/s prefill (2×B70), both above
Ornith-1.5-35B-A3B-DFlash2MXFP4speculative draftfor Ornith-1.5-35B-A3B
Qwen3.8-27BMXFP4 (two independent conversions), NVFP4, FP8, W4A16, GPTQ-Int4, INT4 (AutoRound), bf16 sourcetargetreference point for the vLLM/OpenVINO int4-ov baseline this project compares against
Qwen3.8-27B + native MTP headMXFP4 or GPTQ-Int4targetmulti-token prediction
Qwen3.8-27B-DFlash2MXFP4speculative draftfor Qwen3.8-27B
Qwen3.6-35B-A3B-GPTQ-Int4GPTQ-Int4target35B MoE
Qwen3.6-35B-A3B-DFlashbf16speculative draftfor its own GPTQ-Int4 target
Qwen3.5-35B-A3B-DFlashbf16speculative draftfor Ornith, which shares its base architecture
Qwen3.8-Flash-NextNVFP4targettoo large for one B70's VRAM; tiered across VRAM, system RAM, and SSD
Muse-Glimmer-30BINT4 (W4A16), GPTQ-INT4, MXFP4targethybrid sliding-window / full attention
K2-Horizon-MoVA-36B-A4BMXFP4 (quantized on load from its bf16 release)target36B MoE, its own architecture rather than a Qwen3.5-MoE derivative
Agnes-3.0-FlashMXFP4target

All seven weight formats (BF16, FP8 E4M3/E5M2, INT8, INT4, MXFP8, MXFP4) run through one decode path shared by every model above and by the host-side tests.

Tensor-parallel and pipeline-parallel both run correctly across two B70s, but they exist for checkpoints whose own footprint doesn't fit one card's VRAM — every model in the table fits a single B70 and runs fastest that way.

Speculative decoding (the MTP and DFlash2 rows above) is currently correct but slower than plain decode on every model it's paired with — functional, not yet a throughput win.

Getting started

The engine and build scripts are in this repository, along with a Dockerfile that builds a minimal runtime image -- just GRIMOIRE's own binaries plus Intel's GPU driver stack, nothing else:

docker build -t grimoire-b70 .
docker run --init --stop-timeout 300 --device /dev/dri/renderDXXX \
    -v /path/to/your/models:/models -p 8000:8000 \
    grimoire-b70 server --model /models/<checkpoint> --proj mxfp4 --port 8000

--init and a generous --stop-timeout are not optional: GPU work in flight when a container is killed without them can wedge the card hard enough to need a power cycle. Replace server with generate -m ... -p "..." -n <tokens> for a one-shot CLI run instead of the HTTP server. Check the project status and the model table above before choosing a model or format.

Development notes

GRIMOIRE is under active development. Model support, build requirements, and performance may change as the engine evolves. Benchmarks should be read with their linked commit notes, which record the model, test conditions, correctness checks, and known trade-offs.

doopeworld/GRIMOIRE

C++

1

341 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

GRIMOIRE: native C++/SYCL LLM inference for Intel Arc Pro B70 — Ornith 1.5 hits 175 tok/s single request, 464 tok/s at concurrency 8 (r/LocalLLM)

https://preview.redd.it/azrcl8q94gth1.png?width=2172&amp;format=png&amp;auto=webp&amp;s=20d4010874d090d1d35ee09a292e77696e4820b0 Hi everyone, I’ve been building **GRIMOIRE**, an experimental LLM inference engine for **Intel Battlemage GPUs**, with the **Arc Pro B70** as its primary target. GRIMOIRE…

2

Oct 4, 2026

README

GRIMOIRE

GRIMOIRE LLM inference banner

LLM inference built for Intel Arc Pro B70 Battlemage GPUs.

GRIMOIRE is a native inference engine written in C++ with SYCL and Level Zero. It is being developed to run and optimize large language models directly on Intel Battlemage hardware, with a focus on the Arc Pro B70.

The aim is straightforward: make high-performance local LLM inference possible on this hardware, using kernels and memory paths designed for the GPU instead of relying on a general-purpose inference stack.

Project status: Experimental. GRIMOIRE is actively developed, but it is not yet a stable, production-ready release. Expect incomplete features, changing model support, and performance or compatibility issues. A ready-to-use inference image or downloadable model package will be added here later.

Why GRIMOIRE was created

Most inference tooling and performance guidance is centered on other GPU platforms. Battlemage owners need an engine that treats Intel hardware as a first-class target and makes its real performance measurable.

GRIMOIRE began as a ground-up inference project for the Arc Pro B70. Building the engine directly exposed important details that a port or a high-level wrapper could hide: how the GPU represents quantized weights, where attention kernels can go wrong, how to keep decode bound by useful memory traffic, and what multi-GPU communication actually costs on this setup.

The project exists to turn those findings into working inference code, reproducible checks, and practical performance improvements for Battlemage users.

What it is for

  • Running supported large language models locally on Intel Arc Pro B70 GPUs.
  • Exploring native C++ / SYCL / Level Zero kernels for model inference.
  • Measuring prefill and token-generation performance on real hardware.
  • Improving single-GPU execution and developing multi-GPU paths for models that need more memory.
  • Checking changes against a growing set of model and checkpoint combinations.

GRIMOIRE is a hardware-focused inference engine, not a hosted model service. Model compatibility and performance depend on the architecture, weight format, configuration, and hardware available.

How it is built

  • C++ inference engine
  • SYCL GPU programming
  • Level Zero device runtime
  • Intel Battlemage, with the Arc Pro B70 as the primary target
  • Custom GPU work for operations such as attention, matrix multiplication, quantized weight handling, and mixture-of-experts layers

The project aims to keep inference native to the target hardware and does not use vLLM, PyTorch, or OpenVINO as its inference backend.

Recent measured results

These are results recorded in the linked project commits on the developer's Arc Pro B70 system. They are examples from specific models and test runs, not universal performance guarantees.

WorkloadRecorded resultContext
Ornith-1.5-35B-A3B token generation198.6 tokens/sSingle B70; 256 generated tokens in the recorded run. Commit
Ornith-1.5-35B-A3B prefill10,030–10,165 tokens/sTwo-GPU pipeline-parallel run, 5,987-token prompt. Output matched the single-B70 run for the checked text. Commit
TP decode communication batching59.3–60.0 tokens/sRecorded TP run after combining independent projection gathers; see the commit for the setup and comparison. Commit

The two-GPU figures demonstrate an actively developed path. Tensor and pipeline parallelism are intended for models that need multiple GPUs for capacity; they are not automatically faster for a model that already fits on one card.

Results change as kernels and test coverage evolve. The commit links include the measurements, validation notes, and limitations for each change. The latest model-zoo audit reports 27 confirmed working model/checkpoint combinations in the local regression sweep, with additional formats and configurations correctly identified as out of scope. Latest audit

Supported models

Everything below is validated end-to-end — loaded, generating coherent text, checked against a reference where one exists — on one Intel Arc Pro B70, by the project's own regression suite. That suite runs after every change that touches shared decode, prefill, attention, or MoE code, not as a one-time check.

ModelFormatsRoleNotes
Ornith-1.5-35B-A3BMXFP4, NVFP4, FP8, GPTQ-Int4, INT4 (AutoRound W4A16), bf16 sourcetarget35B MoE, 256 experts / top-8. 198.6 tok/s decode, 10,030–10,165 tok/s prefill (2×B70), both above
Ornith-1.5-35B-A3B-DFlash2MXFP4speculative draftfor Ornith-1.5-35B-A3B
Qwen3.8-27BMXFP4 (two independent conversions), NVFP4, FP8, W4A16, GPTQ-Int4, INT4 (AutoRound), bf16 sourcetargetreference point for the vLLM/OpenVINO int4-ov baseline this project compares against
Qwen3.8-27B + native MTP headMXFP4 or GPTQ-Int4targetmulti-token prediction
Qwen3.8-27B-DFlash2MXFP4speculative draftfor Qwen3.8-27B
Qwen3.6-35B-A3B-GPTQ-Int4GPTQ-Int4target35B MoE
Qwen3.6-35B-A3B-DFlashbf16speculative draftfor its own GPTQ-Int4 target
Qwen3.5-35B-A3B-DFlashbf16speculative draftfor Ornith, which shares its base architecture
Qwen3.8-Flash-NextNVFP4targettoo large for one B70's VRAM; tiered across VRAM, system RAM, and SSD
Muse-Glimmer-30BINT4 (W4A16), GPTQ-INT4, MXFP4targethybrid sliding-window / full attention
K2-Horizon-MoVA-36B-A4BMXFP4 (quantized on load from its bf16 release)target36B MoE, its own architecture rather than a Qwen3.5-MoE derivative
Agnes-3.0-FlashMXFP4target

All seven weight formats (BF16, FP8 E4M3/E5M2, INT8, INT4, MXFP8, MXFP4) run through one decode path shared by every model above and by the host-side tests.

Tensor-parallel and pipeline-parallel both run correctly across two B70s, but they exist for checkpoints whose own footprint doesn't fit one card's VRAM — every model in the table fits a single B70 and runs fastest that way.

Speculative decoding (the MTP and DFlash2 rows above) is currently correct but slower than plain decode on every model it's paired with — functional, not yet a throughput win.

Getting started

The engine and build scripts are in this repository, along with a Dockerfile that builds a minimal runtime image -- just GRIMOIRE's own binaries plus Intel's GPU driver stack, nothing else:

docker build -t grimoire-b70 .
docker run --init --stop-timeout 300 --device /dev/dri/renderDXXX \
    -v /path/to/your/models:/models -p 8000:8000 \
    grimoire-b70 server --model /models/<checkpoint> --proj mxfp4 --port 8000

--init and a generous --stop-timeout are not optional: GPU work in flight when a container is killed without them can wedge the card hard enough to need a power cycle. Replace server with generate -m ... -p "..." -n <tokens> for a one-shot CLI run instead of the HTTP server. Check the project status and the model table above before choosing a model or format.

Development notes

GRIMOIRE is under active development. Model support, build requirements, and performance may change as the engine evolves. Benchmarks should be read with their linked commit notes, which record the model, test conditions, correctness checks, and known trade-offs.

Languages

C++

80.8%

Python

15.4%

Shell

3.4%