An MLIR-based compiler framework bridges DSLs (domain-specific languages) to DSAs (domain-specific architectures).
752
stars
1,500
commits
Python
primary language
Sep 7, 2026
updated
An MLIR-based compiler framework designed for a co-design ecosystem from DSL (domain-specific languages) to DSA (domain-specific architectures). (Project page)
Please make sure the dependencies are available on your machine.
sudo apt install flatbuffers-compiler libflatbuffers-dev libnuma-dev
git clone git@github.com:buddy-compiler/buddy-mlir.git
cd buddy-mlir
git submodule update --init llvm
pip
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uv
uv venv
source .venv/bin/activate
uv pip install -r requirements.txt
conda
conda activate <your virtual environment name>
cd buddy-mlir
pip install -r requirements.txt
cd buddy-mlir
cmake -G Ninja -S llvm/llvm -B llvm/build \
-DLLVM_ENABLE_PROJECTS="mlir;clang" \
-DLLVM_ENABLE_RUNTIMES="openmp" \
-DLLVM_TARGETS_TO_BUILD="host;RISCV" \
-DLLVM_ENABLE_ASSERTIONS=ON \
-DOPENMP_ENABLE_LIBOMPTARGET=OFF \
-DCMAKE_BUILD_TYPE=RELEASE \
-DMLIR_ENABLE_BINDINGS_PYTHON=ON \
-DPython3_EXECUTABLE="$(which python)" \
-DPython_EXECUTABLE="$(which python)"
ninja -C llvm/build check-clang check-mlir check-openmp
If your target machine includes an NVIDIA GPU, you can add the following configuration:
-DLLVM_TARGETS_TO_BUILD="host;RISCV;NVPTX" \
-DMLIR_ENABLE_CUDA_RUNNER=ON \
cd buddy-mlir
cmake -G Ninja -S . -B build \
-DMLIR_DIR=$PWD/llvm/build/lib/cmake/mlir \
-DLLVM_DIR=$PWD/llvm/build/lib/cmake/llvm \
-DLLVM_ENABLE_ASSERTIONS=ON \
-DCMAKE_BUILD_TYPE=RELEASE \
-DBUDDY_MLIR_ENABLE_PYTHON_PACKAGES=ON \
-DPython3_EXECUTABLE="$(which python)" \
-DPython_EXECUTABLE="$(which python)"
ninja -C build
ninja -C build check-buddy
Set the PYTHONPATH environment variable to include both the LLVM/MLIR Python bindings and buddy-mlir Python packages:
export BUDDY_MLIR_BUILD_DIR=$PWD/build
export LLVM_MLIR_BUILD_DIR=$PWD/llvm/build
export PYTHONPATH=${BUDDY_MLIR_BUILD_DIR}/python_packages:${PYTHONPATH}
If you want to test your model end-to-end conversion and inference, you can add the following configuration
cmake -G Ninja -S . -B build -DBUDDY_ENABLE_E2E_TESTS=ON
ninja -C build check-e2e
Use the following to build:
cd buddy-mlir
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32.json \
--build-dir build
For Whisper, use the same build entry point with the Whisper spec:
python3 tools/buddy-codegen/build_model.py \
--spec models/whisper/specs/base.json \
--build-dir build
DeepSeek R1, Whisper, and Qwen3-VL support template-based layer-partitioned compilation. It is disabled by default. Enable it explicitly by passing
--cmake-args=-DBUDDY_MODEL_LAYER_PARTITION=ON to build_model.py. DeepSeek R1 also retains the existing PartitionedGraphDriver workflow. See Layer Partitioning for details.
For example, enable template-based layer partitioning for Whisper with:
python3 tools/buddy-codegen/build_model.py \
--spec models/whisper/specs/base.json \
--build-dir build \
--cmake-args=-DBUDDY_MODEL_LAYER_PARTITION=ON
To import weights from a local HuggingFace style directory (offline or a custom path), pass --local-model to that directory (it must contain config.json and the weight files). If you omit --hf-config, build_model.py uses <local-model>/config.json for codegen when present:
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32.json \
--build-dir build \
--local-model /path/to/DeepSeek-R1-Distill-Qwen-1.5B
If CMake is configured with -DBUDDY_BUILD_DEEPSEEK_R1_MODEL=ON, you can build the model with:
ninja deepseek_r1_model_so deepseek_r1_rax
To build the DeepSeek R1 f32 tiered KV cache variant for use with buddy-cli,
use the dedicated spec:
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32_tiered_kv_cache.json \
--build-dir build
./build/bin/buddy-cli \
--model ./build/models/deepseek_r1/deepseek_r1.rax \
--prompt "Tell me a joke in 200 words."
# Equivalent to: numactl --cpunodebind=0,1,2,3 --interleave=0,1,2,3 taskset -c 0-47
./build/bin/buddy-cli \
--numa 0,1,2,3 \
--cpus 0-47 \
--model ./build/models/deepseek_r1/deepseek_r1.rax \
--prompt "Tell me a joke in 200 words."
Whisper uses the same .rax / buddy-cli deployment path, with an audio input:
./build/bin/buddy-cli \
--model ./build/models/whisper/whisper.rax \
--audio ./build/models/whisper/audio.wav
models/qwen3_vl is a self-contained vision-language model (ViT + DeepStack
encoder feeding a dense Qwen3 decoder) that runs end-to-end on buddy-compiled
kernels via buddy-cli. Use the same tools/buddy-codegen/build_model.py entry
point with the Qwen3-VL spec (a local HuggingFace snapshot is required):
python3 tools/buddy-codegen/build_model.py \
--spec models/qwen3_vl/specs/instruct_2b.json \
--build-dir build \
--local-model /path/to/Qwen3-VL-2B-Instruct
./build/bin/buddy-cli \
--model ./build/models/qwen3_vl/qwen3_vl.rax \
--image ./models/qwen3_vl/test_text.png \
--prompt "Read all the text in the image."
See models/qwen3_vl/README.md for prerequisites and
details.
We use setuptools to bundle CMake outputs (Python packages, bin/, and
lib/) into a single wheel.
Build x86_64 artifacts:
./scripts/release.sh cp312 0.0.0 x86_64
Build riscv64 artifacts:
./scripts/release.sh cp312 0.0.0 riscv64
This script calls docker run internally to enter the offical manylinux container,
builds LLVM and buddy_mlir, and writes artifacts to:
./build-docker/x86_64/<py_tag>/target./build-docker/riscv64/<py_tag>/targetSee Manylinux release notes for current known build notes.
Install and test the wheel:
pip install buddy-*.whl --no-deps
python -c "import buddy; import buddy_mlir; print('ok')"
buddy-opt --help
We provide examples to demonstrate how to use the passes and interfaces in buddy-mlir, including IR-level transformations, domain-specific applications, and testing demonstrations.
For more details, please see the examples documentation.
We welcome contributions to our open-source project!
Before contributing, please read the Contributor Guide and Code Style.
To maintain code quality, this project provides pre-commit checks:
pre-commit install
If you find our project and research useful or refer to it in your own work, please cite the survey paper in which the Buddy Compiler design was first proposed:
@article{zhang2023compiler,
title={Compiler Technologies in Deep Learning Co-Design: A Survey},
author={Zhang, Hongbin and Xing, Mingjie and Wu, Yanjun and Zhao, Chen},
journal={Intelligent Computing},
year={2023},
publisher={AAAS}
}
For direct access to the paper, please visit Compiler Technologies in Deep Learning Co-Design: A Survey.
(top 30 of 80)
Python
47.7%
C++
44.2%
MLIR
3.8%
CMake
3.0%
An MLIR-based compiler framework bridges DSLs (domain-specific languages) to DSAs (domain-specific architectures).
752
stars
1,500
commits
Python
primary language
Sep 7, 2026
updated
An MLIR-based compiler framework designed for a co-design ecosystem from DSL (domain-specific languages) to DSA (domain-specific architectures). (Project page)
Please make sure the dependencies are available on your machine.
sudo apt install flatbuffers-compiler libflatbuffers-dev libnuma-dev
git clone git@github.com:buddy-compiler/buddy-mlir.git
cd buddy-mlir
git submodule update --init llvm
pip
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uv
uv venv
source .venv/bin/activate
uv pip install -r requirements.txt
conda
conda activate <your virtual environment name>
cd buddy-mlir
pip install -r requirements.txt
cd buddy-mlir
cmake -G Ninja -S llvm/llvm -B llvm/build \
-DLLVM_ENABLE_PROJECTS="mlir;clang" \
-DLLVM_ENABLE_RUNTIMES="openmp" \
-DLLVM_TARGETS_TO_BUILD="host;RISCV" \
-DLLVM_ENABLE_ASSERTIONS=ON \
-DOPENMP_ENABLE_LIBOMPTARGET=OFF \
-DCMAKE_BUILD_TYPE=RELEASE \
-DMLIR_ENABLE_BINDINGS_PYTHON=ON \
-DPython3_EXECUTABLE="$(which python)" \
-DPython_EXECUTABLE="$(which python)"
ninja -C llvm/build check-clang check-mlir check-openmp
If your target machine includes an NVIDIA GPU, you can add the following configuration:
-DLLVM_TARGETS_TO_BUILD="host;RISCV;NVPTX" \
-DMLIR_ENABLE_CUDA_RUNNER=ON \
cd buddy-mlir
cmake -G Ninja -S . -B build \
-DMLIR_DIR=$PWD/llvm/build/lib/cmake/mlir \
-DLLVM_DIR=$PWD/llvm/build/lib/cmake/llvm \
-DLLVM_ENABLE_ASSERTIONS=ON \
-DCMAKE_BUILD_TYPE=RELEASE \
-DBUDDY_MLIR_ENABLE_PYTHON_PACKAGES=ON \
-DPython3_EXECUTABLE="$(which python)" \
-DPython_EXECUTABLE="$(which python)"
ninja -C build
ninja -C build check-buddy
Set the PYTHONPATH environment variable to include both the LLVM/MLIR Python bindings and buddy-mlir Python packages:
export BUDDY_MLIR_BUILD_DIR=$PWD/build
export LLVM_MLIR_BUILD_DIR=$PWD/llvm/build
export PYTHONPATH=${BUDDY_MLIR_BUILD_DIR}/python_packages:${PYTHONPATH}
If you want to test your model end-to-end conversion and inference, you can add the following configuration
cmake -G Ninja -S . -B build -DBUDDY_ENABLE_E2E_TESTS=ON
ninja -C build check-e2e
Use the following to build:
cd buddy-mlir
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32.json \
--build-dir build
For Whisper, use the same build entry point with the Whisper spec:
python3 tools/buddy-codegen/build_model.py \
--spec models/whisper/specs/base.json \
--build-dir build
DeepSeek R1, Whisper, and Qwen3-VL support template-based layer-partitioned compilation. It is disabled by default. Enable it explicitly by passing
--cmake-args=-DBUDDY_MODEL_LAYER_PARTITION=ON to build_model.py. DeepSeek R1 also retains the existing PartitionedGraphDriver workflow. See Layer Partitioning for details.
For example, enable template-based layer partitioning for Whisper with:
python3 tools/buddy-codegen/build_model.py \
--spec models/whisper/specs/base.json \
--build-dir build \
--cmake-args=-DBUDDY_MODEL_LAYER_PARTITION=ON
To import weights from a local HuggingFace style directory (offline or a custom path), pass --local-model to that directory (it must contain config.json and the weight files). If you omit --hf-config, build_model.py uses <local-model>/config.json for codegen when present:
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32.json \
--build-dir build \
--local-model /path/to/DeepSeek-R1-Distill-Qwen-1.5B
If CMake is configured with -DBUDDY_BUILD_DEEPSEEK_R1_MODEL=ON, you can build the model with:
ninja deepseek_r1_model_so deepseek_r1_rax
To build the DeepSeek R1 f32 tiered KV cache variant for use with buddy-cli,
use the dedicated spec:
python3 tools/buddy-codegen/build_model.py \
--spec models/deepseek_r1/specs/f32_tiered_kv_cache.json \
--build-dir build
./build/bin/buddy-cli \
--model ./build/models/deepseek_r1/deepseek_r1.rax \
--prompt "Tell me a joke in 200 words."
# Equivalent to: numactl --cpunodebind=0,1,2,3 --interleave=0,1,2,3 taskset -c 0-47
./build/bin/buddy-cli \
--numa 0,1,2,3 \
--cpus 0-47 \
--model ./build/models/deepseek_r1/deepseek_r1.rax \
--prompt "Tell me a joke in 200 words."
Whisper uses the same .rax / buddy-cli deployment path, with an audio input:
./build/bin/buddy-cli \
--model ./build/models/whisper/whisper.rax \
--audio ./build/models/whisper/audio.wav
models/qwen3_vl is a self-contained vision-language model (ViT + DeepStack
encoder feeding a dense Qwen3 decoder) that runs end-to-end on buddy-compiled
kernels via buddy-cli. Use the same tools/buddy-codegen/build_model.py entry
point with the Qwen3-VL spec (a local HuggingFace snapshot is required):
python3 tools/buddy-codegen/build_model.py \
--spec models/qwen3_vl/specs/instruct_2b.json \
--build-dir build \
--local-model /path/to/Qwen3-VL-2B-Instruct
./build/bin/buddy-cli \
--model ./build/models/qwen3_vl/qwen3_vl.rax \
--image ./models/qwen3_vl/test_text.png \
--prompt "Read all the text in the image."
See models/qwen3_vl/README.md for prerequisites and
details.
We use setuptools to bundle CMake outputs (Python packages, bin/, and
lib/) into a single wheel.
Build x86_64 artifacts:
./scripts/release.sh cp312 0.0.0 x86_64
Build riscv64 artifacts:
./scripts/release.sh cp312 0.0.0 riscv64
This script calls docker run internally to enter the offical manylinux container,
builds LLVM and buddy_mlir, and writes artifacts to:
./build-docker/x86_64/<py_tag>/target./build-docker/riscv64/<py_tag>/targetSee Manylinux release notes for current known build notes.
Install and test the wheel:
pip install buddy-*.whl --no-deps
python -c "import buddy; import buddy_mlir; print('ok')"
buddy-opt --help
We provide examples to demonstrate how to use the passes and interfaces in buddy-mlir, including IR-level transformations, domain-specific applications, and testing demonstrations.
For more details, please see the examples documentation.
We welcome contributions to our open-source project!
Before contributing, please read the Contributor Guide and Code Style.
To maintain code quality, this project provides pre-commit checks:
pre-commit install
If you find our project and research useful or refer to it in your own work, please cite the survey paper in which the Buddy Compiler design was first proposed:
@article{zhang2023compiler,
title={Compiler Technologies in Deep Learning Co-Design: A Survey},
author={Zhang, Hongbin and Xing, Mingjie and Wu, Yanjun and Zhao, Chen},
journal={Intelligent Computing},
year={2023},
publisher={AAAS}
}
For direct access to the paper, please visit Compiler Technologies in Deep Learning Co-Design: A Survey.
(top 30 of 80)
Python
47.7%
C++
44.2%
MLIR
3.8%
CMake
3.0%