[MobiSys 2026] On-device ultra-low-bit LLM inference with LUT.
C++
27
7 commits
updated Jun 24, 2026
vlut.cpp (lookup table-based) vs. llama.cpp (dequantization-based) running Llama3-8B-1.58-100B-tokens on Intel Core Ultra 7 258V (see run_batched_decode.sh):
https://github.com/user-attachments/assets/d38908fb-a5f4-4a48-ba44-373057c3de39
vlut.cpp vs. llama.cpp vs. T-MAC in GeMM kernel benchmark (see Evaluation.md for a detailed evaluation guide):

vlut.cpp is a lightweight extension of llama.cpp that implements Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices. It targets parallel ultra-low-bit LLM inference. Parallel scenarios include:
The Vec-LUT kernel is fast with:
Based on the Vec-LUT kernel, vlut.cpp is efficient and easy to use with:
vlut.cpp supports all mainstream CPUs (Intel, AMD, ARM), and operating systems (Linux, Android, Mac OS, Windows). You can build and test vlut.cpp on almost any platforms.
We recommend using the Windows Subsystem for Linux (WSL) on Windows, and Termux on Android. They provide Linux-like development environments.
Please refer to Evaluation.md for recommended specifications to run the evaluation.
vlut.cpp now supports a rich set of ternary (1.58-bit) LLMs:
1bitLLM/bitnet_b1_58-3BHF1BitLLM/Llama3-8B-1.58-100B-tokenstiiuae/Falcon3-1B-Instruct-1.58bitSpectraSuite/TriLM_3.9B_UnpackedWe provide pre-converted models for immediate on-device deployment with vlut.cpp. Check out 1.58-bit LLMs for vlut.cpp on Huggingface.
This section walks you through the minimum steps required to run a ternary LLM with vlut.cpp:
llama-batched or benchmark with llama-bench.For a more detailed evaluation pipeline (GeMM, prefill, batched decoding, multi-framework comparison), see Evaluation.md.
vlut.cpp follows the same build process as llama.cpp (CPU build), see how to build.
Run the following commands to build vlut.cpp with 4 parallel jobs:
cmake -B build
cmake --build build --config Release -j 4
Before quantization, HuggingFace models (safetensors) must be converted to vlut GGUF. Skip this step if you use our pre-converted models.
Install dependencies:
pip install -r requirements.txt
Convert a model (BitNet 3B for example):
python ./convert_hf_to_gguf_vlut.py ~/models/bitnet_b1_58-3B --outfile ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf
vlut.cpp provides lossless ternary packings I1 and I2, with optional K-tiling variants (e.g., I1_V_2, I2_V_4). Skip this step if you use our pre-converted models.
Quantize the converted GGUF:
./build/bin/llama-quantize ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf I1_V_2
./build/bin/llama-quantize ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf I2_V_4
The quantized model will be saved as ggml-model-{quant_type}.gguf.
Use llama-batched to run a parallel inference:
./build/bin/llama-batched \
-m ~/models/bitnet_b1_58-3B/ggml-model-I2_V_4.gguf \
-p "I believe the meaning of life is" \
-np 32 -n 16 -t 1 --temp 0.5 --repeat-penalty 1.5
llama-bench lets you measure the performance of the inference for various parameters.
Example:
./build/bin/llama-bench -m model.gguf -t 4 -p 128 -n 0
This project is built atop llama.cpp. Thanks to all the contributors for their valuable works!
The LUT-based idea is inspired by T-MAC, which is primarily optimized for non-parallel scenarios (e.g., single-batch decoding).
If you find this project useful, please cite our paper:
@inproceedings{li2026vec,
title={Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices},
author={Li, Xiangyu and Yin, Chengyu and Wang, Weijun and Wei, Jianyu and Cao, Ting and Liu, Yunxin},
booktitle={Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services},
pages={216--230},
year={2026}
}
C++
58.6%
C
18.0%
Python
9.9%
Cuda
6.4%
Metal
2.4%
Objective-C
2.4%
CMake
1.2%
[MobiSys 2026] On-device ultra-low-bit LLM inference with LUT.
C++
27
7 commits
updated Jun 24, 2026
vlut.cpp (lookup table-based) vs. llama.cpp (dequantization-based) running Llama3-8B-1.58-100B-tokens on Intel Core Ultra 7 258V (see run_batched_decode.sh):
https://github.com/user-attachments/assets/d38908fb-a5f4-4a48-ba44-373057c3de39
vlut.cpp vs. llama.cpp vs. T-MAC in GeMM kernel benchmark (see Evaluation.md for a detailed evaluation guide):

vlut.cpp is a lightweight extension of llama.cpp that implements Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices. It targets parallel ultra-low-bit LLM inference. Parallel scenarios include:
The Vec-LUT kernel is fast with:
Based on the Vec-LUT kernel, vlut.cpp is efficient and easy to use with:
vlut.cpp supports all mainstream CPUs (Intel, AMD, ARM), and operating systems (Linux, Android, Mac OS, Windows). You can build and test vlut.cpp on almost any platforms.
We recommend using the Windows Subsystem for Linux (WSL) on Windows, and Termux on Android. They provide Linux-like development environments.
Please refer to Evaluation.md for recommended specifications to run the evaluation.
vlut.cpp now supports a rich set of ternary (1.58-bit) LLMs:
1bitLLM/bitnet_b1_58-3BHF1BitLLM/Llama3-8B-1.58-100B-tokenstiiuae/Falcon3-1B-Instruct-1.58bitSpectraSuite/TriLM_3.9B_UnpackedWe provide pre-converted models for immediate on-device deployment with vlut.cpp. Check out 1.58-bit LLMs for vlut.cpp on Huggingface.
This section walks you through the minimum steps required to run a ternary LLM with vlut.cpp:
llama-batched or benchmark with llama-bench.For a more detailed evaluation pipeline (GeMM, prefill, batched decoding, multi-framework comparison), see Evaluation.md.
vlut.cpp follows the same build process as llama.cpp (CPU build), see how to build.
Run the following commands to build vlut.cpp with 4 parallel jobs:
cmake -B build
cmake --build build --config Release -j 4
Before quantization, HuggingFace models (safetensors) must be converted to vlut GGUF. Skip this step if you use our pre-converted models.
Install dependencies:
pip install -r requirements.txt
Convert a model (BitNet 3B for example):
python ./convert_hf_to_gguf_vlut.py ~/models/bitnet_b1_58-3B --outfile ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf
vlut.cpp provides lossless ternary packings I1 and I2, with optional K-tiling variants (e.g., I1_V_2, I2_V_4). Skip this step if you use our pre-converted models.
Quantize the converted GGUF:
./build/bin/llama-quantize ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf I1_V_2
./build/bin/llama-quantize ~/models/bitnet_b1_58-3B/bitnet_b1_58-3B.vlut.gguf I2_V_4
The quantized model will be saved as ggml-model-{quant_type}.gguf.
Use llama-batched to run a parallel inference:
./build/bin/llama-batched \
-m ~/models/bitnet_b1_58-3B/ggml-model-I2_V_4.gguf \
-p "I believe the meaning of life is" \
-np 32 -n 16 -t 1 --temp 0.5 --repeat-penalty 1.5
llama-bench lets you measure the performance of the inference for various parameters.
Example:
./build/bin/llama-bench -m model.gguf -t 4 -p 128 -n 0
This project is built atop llama.cpp. Thanks to all the contributors for their valuable works!
The LUT-based idea is inspired by T-MAC, which is primarily optimized for non-parallel scenarios (e.g., single-batch decoding).
If you find this project useful, please cite our paper:
@inproceedings{li2026vec,
title={Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices},
author={Li, Xiangyu and Yin, Chengyu and Wang, Weijun and Wei, Jianyu and Cao, Ting and Liu, Yunxin},
booktitle={Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services},
pages={216--230},
year={2026}
}
C++
58.6%
C
18.0%
Python
9.9%
Cuda
6.4%
Metal
2.4%
Objective-C
2.4%
CMake
1.2%