ciresnave/fuel

Rust ML Framework

Rust

0

5,374 commits

updated Sep 17, 2026

See the code

README

fuel

discord server License License

Fuel is a minimalist ML framework for Rust with a focus on performance (including GPU support) and ease of use. Try our online demos: whisper, LLaMA2, T5, yolo, Segment Anything.

Architecture

For the durable description of what fuel is and how it's structured, see docs/architecture/. The architecture set covers the IR, the optimization model, the backend contract, runtime, tolerance, persistence, and what fuel deliberately doesn't try to be. For the in-flight phase work and the planned order, see ROADMAP.md.

Get started

Make sure that you have fuel-core correctly installed as described in Installation.

Let's see how to run a simple matrix multiplication. Write the following to your myapp/src/main.rs file:

use fuel_core::{Device, Tensor};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let device = Device::Cpu;

    let a = Tensor::randn(0f32, 1., (2, 3), &device)?;
    let b = Tensor::randn(0f32, 1., (3, 4), &device)?;

    let c = a.matmul(&b)?;
    println!("{c}");
    Ok(())
}

cargo run should display a tensor of shape Tensor[[2, 4], f32].

Having installed fuel with Cuda support, simply define the device to be on GPU:

- let device = Device::cpu();
+ let device = fuel_core::cuda_backend::new_device(0)?;

For more advanced examples, please have a look at the following section.

Check out our examples

These online demos run entirely in your browser:

We also provide some command line based examples using state of the art models:

  • LLaMA v1, v2, and v3: general LLM, includes the SOLAR-10.7B variant.
  • Falcon: general LLM.
  • Codegeex4: Code completion, code interpreter, web search, function calling, repository-level
  • GLM4: Open Multilingual Multimodal Chat LMs by THUDM
  • Gemma v1 and v2: 2b and 7b+/9b general LLMs from Google Deepmind.
  • RecurrentGemma: 2b and 7b Griffin based models from Google that mix attention with a RNN like state.
  • Phi-1, Phi-1.5, Phi-2, and Phi-3: 1.3b, 2.7b, and 3.8b general LLMs with performance on par with 7b models.
  • StableLM-3B-4E1T: a 3b general LLM pre-trained on 1T tokens of English and code datasets. Also supports StableLM-2, a 1.6b LLM trained on 2T tokens, as well as the code variants.
  • Mamba: an inference only implementation of the Mamba state space model.
  • Mistral7b-v0.1: a 7b general LLM with better performance than all publicly available 13b models as of 2023-09-28.
  • Mixtral8x7b-v0.1: a sparse mixture of experts 8x7b general LLM with better performance than a Llama 2 70B model with much faster inference.
  • StarCoder and StarCoder2: LLM specialized to code generation.
  • Qwen1.5: Bilingual (English/Chinese) LLMs.
  • RWKV v5 and v6: An RNN with transformer level LLM performance.
  • Replit-code-v1.5: a 3.3b LLM specialized for code completion.
  • Yi-6B / Yi-34B: two bilingual (English/Chinese) general LLMs with 6b and 34b parameters.
  • Quantized LLaMA: quantized version of the LLaMA model using the same quantization techniques as llama.cpp.
  • Quantized Qwen3 MoE: support gguf quantized models of Qwen3 MoE models.
  • Stable Diffusion: text to image generative model, support for the 1.5, 2.1, SDXL 1.0 and Turbo versions.
  • Wuerstchen: another text to image generative model.

  • SegFormer: transformer based semantic segmentation model.
  • Whisper: speech recognition model.
  • EnCodec: high-quality audio compression model using residual vector quantization.
  • MetaVoice: foundational model for text-to-speech.
  • Parler-TTS: large text-to-speech model.
  • T5, Bert, JinaBert : useful for sentence embeddings.
  • DINOv2: computer vision model trained using self-supervision (can be used for imagenet classification, depth evaluation, segmentation).
  • VGG, RepVGG: computer vision models.
  • BLIP: image to text model, can be used to generate captions for an image.
  • CLIP: multi-model vision and language model.
  • TrOCR: a transformer OCR model, with dedicated submodels for hand-writing and printed recognition.
  • Marian-MT: neural machine translation model, generates the translated text from the input text.
  • Moondream: tiny computer-vision model that can answer real-world questions about images.

Run them using commands like:

cargo run --example quantized --release

In order to use CUDA add --features cuda to the example command line. If you have cuDNN installed, use --features cudnn for even more speedups.

Cargo feature flags

Fuel is designed so that a CPU-only build compiles without any GPU toolkit installed. GPU support is opt-in via Cargo feature flags:

FeatureWhat it enablesRequires
(none)CPU-only build (portable Rust gemm). No GPU code compiled.
cudaNVIDIA GPU backend (cuBLAS, cuDNN).CUDA toolkit ≥ 11
cudnnEnables cuda + cuDNN accelerated conv/norm ops.CUDA toolkit + cuDNN
ncclMulti-GPU communication via NVIDIA NCCL.CUDA + NCCL runtime
vulkanCross-vendor GPU backend via Vulkan (precompiled SPIR-V).Vulkan ≥ 1.3 loader
metalApple Silicon / macOS GPU backend (Metal).macOS 13+
accelerateApple Accelerate BLAS (CPU, macOS only).macOS
mklIntel MKL BLAS (CPU, Linux/Windows). Faster on Intel CPUs.Intel oneMKL runtime
aoclAMD AOCL-BLAS / BLIS (CPU). Faster on Zen-class AMD CPUs.AMD AOCL runtime

Multiple CPU backends can coexist: with --features mkl,aocl both will register on startup, the Phase 6b judge profiles each, and the dispatch table picks the winner per (op, dtype, size_class) empirically. The "wrong" backend for a given CPU (MKL on AMD, AOCL on Intel) just loses the profile race — it doesn't break anything, and there's no need to gate via #[cfg(target_arch)] heuristics.

Runtime requirements (where the shared libraries come from)

Cargo features enable the Rust glue. The actual numerical kernels live in vendor-shipped shared libraries that must be resolvable by the OS dynamic loader at runtime — not at compile time. cargo build --features X succeeds without the runtime present; the binary fails on first call when it can't find the DLL / .so / .dylib.

The required runtime library, default install path, and how the OS loader finds it for each feature:

cudanvcuda.dll / libcuda.so from the NVIDIA driver install. The driver installer adds it to system PATH / ld.so config automatically.

cudnncudnn*.dll / libcudnn.so from the CUDA toolkit or a standalone cuDNN install. Add <cuda>/bin to PATH (Windows) or <cuda>/lib64 to LD_LIBRARY_PATH (Linux).

vulkanvulkan-1.dll / libvulkan.so.1 from the GPU driver install. The driver installer adds it to system PATH / ld.so config automatically.

metal / accelerate — built into macOS. Always available.

mklmkl_rt.2.dll / libmkl_rt.so.2. Default install paths: C:\Program Files (x86)\Intel\oneAPI\mkl\<ver> on Windows; /opt/intel/oneapi/mkl/<ver> on Linux. Run setvars.bat / source setvars.sh from the oneAPI install dir, OR add the MKL redist / lib/intel64 directory to PATH / LD_LIBRARY_PATH.

aoclAOCL-LibBlis-Win-dll.dll / libblis.so. Default install paths: C:\Program Files\AMD\AOCL-Windows\amd-blis\lib\LP64 on Windows; /opt/AMD/aocl-linux-*/aocl-blis/lib/LP64 on Linux. Not added to system PATH by the AOCL installer. Add the lib/LP64 directory above to PATH / LD_LIBRARY_PATH manually before running.

If a runtime library is missing or off the loader's search path, you'll see errors like STATUS_DLL_NOT_FOUND (Windows error 0xc0000135) or error while loading shared libraries (Linux). The fix is always "make the directory containing the named library visible to the dynamic loader" — either through PATH/LD_LIBRARY_PATH, the OS config files (/etc/ld.so.conf.d/ on Linux), or by copying the DLL next to your executable on Windows.

Backends self-test on startup. AoclBackend::try_new() runs a 2×2 sgemm to verify the library actually loaded; if the DLL is missing the call returns Err, the backend doesn't register, and the rest of Fuel transparently falls back to other CPU backends. You won't see a hard crash from a missing optional runtime — only an eprintln! from the probe collector.

On Windows, both AoclBackend::try_new and MklBackend::try_new discover the vendor's BLIS / mkl_rt DLL automatically — they look at standard install paths and the AOCL_ROOT / MKLROOT env vars and prepend the matching bin directory to the process's PATH before the load probe. So cargo run --features aocl,onemkl works out of the box on a normal AOCL / oneAPI install without any manual setvars.bat or path-extension shell prep.

Activating empirical backend selection

Compiling with --features aocl,onemkl registers both backends, but by default Tensor::realize_f32() keeps using the portable Rust gemm — exactly as it did before the per-vendor backends existed. To switch on per-op empirical routing, the app calls populate_dispatch_table() once:

use fuel_core::dispatch;

fn main() -> fuel_core::Result<()> {
    // Option 1: blocking on the main thread. First run measures every
    // backend × op × size_class (~10–60s depending on hardware) and
    // persists the profile to disk. Every subsequent run loads from
    // disk in sub-millisecond.
    dispatch::populate_dispatch_table()?;

    // Option 2: background thread. Routing kicks in once the judge
    // returns; the first few realize calls fall through to the
    // portable CPU baseline, which is fine.
    std::thread::spawn(|| {
        let _ = dispatch::populate_dispatch_table();
    });

    // Option 3: skip the call entirely. realize_f32 keeps using the
    // portable CPU path; no behaviour change. The disk-cache lazy-load
    // means a previous process's `populate_dispatch_table()` is still
    // honored — `dispatch::cached()` quietly loads it on first use.

    // ... your model code uses Tensor::realize_f32() as normal ...
    Ok(())
}

Once a dispatch table is cached, every Tensor::realize_f32() call consults it per op. On a Zen-class AMD CPU with both AOCL and oneMKL enabled, this typically picks AOCL or MKL (whichever wins the empirical race that run) for matmul-heavy work and stays on the portable backend for the few percent that's elementwise. No code changes downstream — the realize_f32() call site is identical to the no-routing default.

If a previous profile becomes stale (driver upgrade, BLAS lib swap, OS kernel update with measurably different behaviour), call dispatch::invalidate(). The next populate_dispatch_table() re-runs the judge and overwrites the persisted profile.

Where to download the vendor runtimes

To add a GPU backend when running examples:

# NVIDIA GPU
cargo run --features cuda --example <name> --release

# Apple Silicon
cargo run --features metal --example <name> --release

# CPU only (no GPU toolkit needed)
cargo run --example <name> --release

To add GPU support to your own project:

# Cargo.toml
[dependencies]
fuel-core = { version = "0.10.2", features = ["cuda"] }   # NVIDIA
fuel-core = { version = "0.10.2", features = ["metal"] }  # Apple
fuel-core = { version = "0.10.2" }                        # CPU only

There are also some wasm examples for whisper and llama2.c. You can either build them with trunk or try them online: whisper, llama2, T5, Phi-1.5, and Phi-2, Segment Anything Model.

For LLaMA2, run the following command to retrieve the weight files and start a test server:

cd fuel-wasm-examples/llama2-c
wget https://huggingface.co/spaces/lmz/candle-llama2/resolve/main/model.bin
wget https://huggingface.co/spaces/lmz/candle-llama2/resolve/main/tokenizer.json
trunk serve --release --port 8081

And then head over to http://localhost:8081/.

Useful External Resources

  • candle-tutorial: A very detailed tutorial showing how to convert a PyTorch model to Fuel.
  • candle-lora: Efficient and ergonomic LoRA implementation for Fuel. fuel-lora has
    out-of-the-box LoRA support for many models from Fuel, which can be found here.
  • candle-video: Rust library for text-to-video generation (LTX-Video and related models) built on Candle, focused on fast, Python-free inference.
  • optimisers: A collection of optimisers including SGD with momentum, AdaGrad, AdaDelta, AdaMax, NAdam, RAdam, and RMSprop.
  • candle-vllm: Efficient platform for inference and serving local LLMs including an OpenAI compatible API server.
  • candle-ext: An extension library to Candle that provides PyTorch functions not currently available in Candle.
  • candle-coursera-ml: Implementation of ML algorithms from Coursera's Machine Learning Specialization course.
  • kalosm: A multi-modal meta-framework in Rust for interfacing with local pre-trained models with support for controlled generation, custom samplers, in-memory vector databases, audio transcription, and more.
  • candle-sampling: Sampling techniques for Candle.
  • gpt-from-scratch-rs: A port of Andrej Karpathy's Let's build GPT tutorial on YouTube showcasing the Fuel API on a toy problem.
  • candle-einops: A pure rust implementation of the python einops library.
  • atoma-infer: A Rust library for fast inference at scale, leveraging FlashAttention2 for efficient attention computation, PagedAttention for efficient KV-cache memory management, and multi-GPU support. It is OpenAI api compatible.
  • llms-from-scratch-rs: A comprehensive Rust translation of the code from Sebastian Raschka's Build an LLM from Scratch book.
  • vllm.rs: A minimalist vLLM implementation in Rust based on Candle.

If you have an addition to this list, please submit a pull request.

Features

  • Simple syntax, looks and feels like PyTorch.
  • Backends.
    • Optimized CPU backend with optional MKL support for x86 and Accelerate for macs.
    • CUDA backend for efficiently running on GPUs, multiple GPU distribution via NCCL.
    • WASM support, run your models in a browser.
  • Included models.
    • Language Models.
      • LLaMA v1, v2, and v3 with variants such as SOLAR-10.7B.
      • Falcon.
      • StarCoder, StarCoder2.
      • Phi 1, 1.5, 2, and 3.
      • Mamba, Minimal Mamba
      • Gemma v1 2b and 7b+, v2 2b and 9b.
      • Mistral 7b v0.1.
      • Mixtral 8x7b v0.1.
      • StableLM-3B-4E1T, StableLM-2-1.6B, Stable-Code-3B.
      • Replit-code-v1.5-3B.
      • Bert.
      • Yi-6B and Yi-34B.
      • Qwen1.5, Qwen1.5 MoE, Qwen3 MoE.
      • RWKV v5 and v6.
    • Quantized LLMs.
      • Llama 7b, 13b, 70b, as well as the chat and code variants.
      • Mistral 7b, and 7b instruct.
      • Mixtral 8x7b.
      • Zephyr 7b a and b (Mistral-7b based).
      • OpenChat 3.5 (Mistral-7b based).
      • Qwen3 MoE (16B-A3B, 32B-A3B)
    • Text to text.
      • T5 and its variants: FlanT5, UL2, MADLAD400 (translation), CoEdit (Grammar correction).
      • Marian MT (Machine Translation).
    • Text to image.
      • Stable Diffusion v1.5, v2.1, XL v1.0.
      • Wurstchen v2.
    • Image to text.
      • BLIP.
      • TrOCR.
    • Audio.
      • Whisper, multi-lingual speech-to-text.
      • EnCodec, audio compression model.
      • MetaVoice-1B, text-to-speech model.
      • Parler-TTS, text-to-speech model.
    • Computer Vision Models.
      • DINOv2, ConvMixer, EfficientNet, ResNet, ViT, VGG, RepVGG, ConvNeXT, ConvNeXTv2, MobileOne, EfficientVit (MSRA), MobileNetv4, Hiera, FastViT.
      • yolo-v3, yolo-v8.
      • Segment-Anything Model (SAM).
      • SegFormer.
  • File formats: load models from safetensors, npz, ggml, or PyTorch files.
  • Serverless (on CPU), small and fast deployments.
  • Quantization support using the llama.cpp quantized types.

How to use

Cheatsheet:

Using PyTorchUsing Fuel
Creationtorch.Tensor([[1, 2], [3, 4]])Tensor::new(&[[1f32, 2.], [3., 4.]], &Device::Cpu)?
Creationtorch.zeros((2, 2))Tensor::zeros((2, 2), DType::F32, &Device::Cpu)?
Indexingtensor[:, :4]tensor.i((.., ..4))?
Operationstensor.view((2, 2))tensor.reshape((2, 2))?
Operationsa.matmul(b)a.matmul(&b)?
Arithmetica + b&a + &b
Devicetensor.to(device="cuda")tensor.to_device(&fuel_core::cuda_backend::new_device(0)?)?
Dtypetensor.to(dtype=torch.float16)tensor.to_dtype(&DType::F16)?
Savingtorch.save({"A": A}, "model.bin")fuel::safetensors::save(&HashMap::from([("A", A)]), "model.safetensors")?
Loadingweights = torch.load("model.bin")fuel::safetensors::load("model.safetensors", &device)

Structure

FAQ

Why should I use Fuel?

Fuel's core goal is to make serverless inference possible. Full machine learning frameworks like PyTorch are very large, which makes creating instances on a cluster slow. Fuel allows deployment of lightweight binaries.

Secondly, Fuel lets you remove Python from production workloads. Python overhead can seriously hurt performance, and the GIL is a notorious source of headaches.

Finally, Rust is cool! A lot of the HF ecosystem already has Rust crates, like safetensors and tokenizers.

Other ML frameworks

  • dfdx is a formidable crate, with shapes being included in types. This prevents a lot of headaches by getting the compiler to complain about shape mismatches right off the bat. However, we found that some features still require nightly, and writing code can be a bit daunting for non rust experts.

    We're leveraging and contributing to other core crates for the runtime so hopefully both crates can benefit from each other.

  • burn is a general crate that can leverage multiple backends so you can choose the best engine for your workload.

  • tch-rs Bindings to the torch library in Rust. Extremely versatile, but they bring in the entire torch library into the runtime. The main contributor of tch-rs is also involved in the development of fuel.

Common Errors

Missing symbols when compiling with the mkl feature.

If you get some missing symbols when compiling binaries/tests using the mkl or accelerate features, e.g. for mkl you get:

  = note: /usr/bin/ld: (....o): in function `blas::sgemm':
          .../blas-0.22.0/src/lib.rs:1944: undefined reference to `sgemm_' collect2: error: ld returned 1 exit status

  = note: some `extern` functions couldn't be found; some native libraries may need to be installed or have their path specified
  = note: use the `-l` flag to specify native libraries to link
  = note: use the `cargo:rustc-link-lib` directive to specify the native libraries to link with Cargo

or for accelerate:

Undefined symbols for architecture arm64:
            "_dgemm_", referenced from:
                fuel_core::accelerate::dgemm::h1b71a038552bcabe in libfuel_core...
            "_sgemm_", referenced from:
                fuel_core::accelerate::sgemm::h2cf21c592cba3c47 in libfuel_core...
          ld: symbol(s) not found for architecture arm64

This is likely due to a missing linker flag that was needed to enable the mkl library. You can try adding the following for mkl at the top of your binary:

extern crate intel_mkl_src;

or for accelerate:

extern crate accelerate_src;

Cannot run the LLaMA examples: access to source requires login credentials

Error: request error: https://huggingface.co/meta-llama/Llama-2-7b-hf/resolve/main/tokenizer.json: status code 401

This is likely because you're not permissioned for the LLaMA-v2 model. To fix this, you have to register on the huggingface-hub, accept the LLaMA-v2 model conditions, and set up your authentication token. See issue #350 for more details.

Docker build

When building CUDA kernels inside a Dockerfile, nvidia-smi cannot be used to auto-detect compute capability.

You must explicitly set CUDA_COMPUTE_CAP, for example:

FROM nvidia/cuda:12.9.0-devel-ubuntu22.04

# Install git and curl
RUN set -eux; \
  apt-get update; \
  apt-get install -y curl git ca-certificates;

# Install Rust
RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y

# Clone fuel repo
RUN git clone https://github.com/ciresnave/fuel.git

# Set compute capability for the build
ARG CUDA_COMPUTE_CAP=90
ENV CUDA_COMPUTE_CAP=${CUDA_COMPUTE_CAP}

# Build with explicit compute cap
WORKDIR /app
COPY . .
RUN cargo build --release features cuda

Compiling with flash-attention fails

/usr/include/c++/11/bits/std_function.h:530:146: error: parameter packs not expanded with ‘...’:

This is a bug in gcc-11 triggered by the Cuda compiler. To fix this, install a different, supported gcc version - for example gcc-10, and specify the path to the compiler in the NVCC_CCBIN environment variable.

env NVCC_CCBIN=/usr/lib/gcc/x86_64-linux-gnu/10 cargo ...

Extremely slow model load time with WSL

This may be caused by the models being loaded from /mnt/c, more details on stackoverflow.

Tracking down errors

You can set RUST_BACKTRACE=1 to be provided with backtraces when a fuel error is generated.

CudaRC error

If you encounter an error like this one called Result::unwrap()on anErr value: LoadLibraryExW { source: Os { code: 126, kind: Uncategorized, message: "The specified module could not be found." } } on windows. To fix copy and rename these 3 files (make sure they are in path). The paths depend on your cuda version. c:\Windows\System32\nvcuda.dll -> cuda.dll c:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4\bin\cublas64_12.dll -> cublas.dll c:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4\bin\curand64_10.dll -> curand.dll

Contributors

(top 30 of 252)

ciresnave

2,635 commits

LaurentMazare

1,603 commits

Narsil

275 commits

ciresnave-bot

166 commits

ciresnave/fuel

Rust ML Framework

Rust

0

5,374 commits

updated Sep 17, 2026

See the code

README

fuel

discord server License License

Fuel is a minimalist ML framework for Rust with a focus on performance (including GPU support) and ease of use. Try our online demos: whisper, LLaMA2, T5, yolo, Segment Anything.

Architecture

For the durable description of what fuel is and how it's structured, see docs/architecture/. The architecture set covers the IR, the optimization model, the backend contract, runtime, tolerance, persistence, and what fuel deliberately doesn't try to be. For the in-flight phase work and the planned order, see ROADMAP.md.

Get started

Make sure that you have fuel-core correctly installed as described in Installation.

Let's see how to run a simple matrix multiplication. Write the following to your myapp/src/main.rs file:

use fuel_core::{Device, Tensor};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let device = Device::Cpu;

    let a = Tensor::randn(0f32, 1., (2, 3), &device)?;
    let b = Tensor::randn(0f32, 1., (3, 4), &device)?;

    let c = a.matmul(&b)?;
    println!("{c}");
    Ok(())
}

cargo run should display a tensor of shape Tensor[[2, 4], f32].

Having installed fuel with Cuda support, simply define the device to be on GPU:

- let device = Device::cpu();
+ let device = fuel_core::cuda_backend::new_device(0)?;

For more advanced examples, please have a look at the following section.

Check out our examples

These online demos run entirely in your browser:

We also provide some command line based examples using state of the art models:

  • LLaMA v1, v2, and v3: general LLM, includes the SOLAR-10.7B variant.
  • Falcon: general LLM.
  • Codegeex4: Code completion, code interpreter, web search, function calling, repository-level
  • GLM4: Open Multilingual Multimodal Chat LMs by THUDM
  • Gemma v1 and v2: 2b and 7b+/9b general LLMs from Google Deepmind.
  • RecurrentGemma: 2b and 7b Griffin based models from Google that mix attention with a RNN like state.
  • Phi-1, Phi-1.5, Phi-2, and Phi-3: 1.3b, 2.7b, and 3.8b general LLMs with performance on par with 7b models.
  • StableLM-3B-4E1T: a 3b general LLM pre-trained on 1T tokens of English and code datasets. Also supports StableLM-2, a 1.6b LLM trained on 2T tokens, as well as the code variants.
  • Mamba: an inference only implementation of the Mamba state space model.
  • Mistral7b-v0.1: a 7b general LLM with better performance than all publicly available 13b models as of 2023-09-28.
  • Mixtral8x7b-v0.1: a sparse mixture of experts 8x7b general LLM with better performance than a Llama 2 70B model with much faster inference.
  • StarCoder and StarCoder2: LLM specialized to code generation.
  • Qwen1.5: Bilingual (English/Chinese) LLMs.
  • RWKV v5 and v6: An RNN with transformer level LLM performance.
  • Replit-code-v1.5: a 3.3b LLM specialized for code completion.
  • Yi-6B / Yi-34B: two bilingual (English/Chinese) general LLMs with 6b and 34b parameters.
  • Quantized LLaMA: quantized version of the LLaMA model using the same quantization techniques as llama.cpp.
  • Quantized Qwen3 MoE: support gguf quantized models of Qwen3 MoE models.
  • Stable Diffusion: text to image generative model, support for the 1.5, 2.1, SDXL 1.0 and Turbo versions.
  • Wuerstchen: another text to image generative model.

  • SegFormer: transformer based semantic segmentation model.
  • Whisper: speech recognition model.
  • EnCodec: high-quality audio compression model using residual vector quantization.
  • MetaVoice: foundational model for text-to-speech.
  • Parler-TTS: large text-to-speech model.
  • T5, Bert, JinaBert : useful for sentence embeddings.
  • DINOv2: computer vision model trained using self-supervision (can be used for imagenet classification, depth evaluation, segmentation).
  • VGG, RepVGG: computer vision models.
  • BLIP: image to text model, can be used to generate captions for an image.
  • CLIP: multi-model vision and language model.
  • TrOCR: a transformer OCR model, with dedicated submodels for hand-writing and printed recognition.
  • Marian-MT: neural machine translation model, generates the translated text from the input text.
  • Moondream: tiny computer-vision model that can answer real-world questions about images.

Run them using commands like:

cargo run --example quantized --release

In order to use CUDA add --features cuda to the example command line. If you have cuDNN installed, use --features cudnn for even more speedups.

Cargo feature flags

Fuel is designed so that a CPU-only build compiles without any GPU toolkit installed. GPU support is opt-in via Cargo feature flags:

FeatureWhat it enablesRequires
(none)CPU-only build (portable Rust gemm). No GPU code compiled.
cudaNVIDIA GPU backend (cuBLAS, cuDNN).CUDA toolkit ≥ 11
cudnnEnables cuda + cuDNN accelerated conv/norm ops.CUDA toolkit + cuDNN
ncclMulti-GPU communication via NVIDIA NCCL.CUDA + NCCL runtime
vulkanCross-vendor GPU backend via Vulkan (precompiled SPIR-V).Vulkan ≥ 1.3 loader
metalApple Silicon / macOS GPU backend (Metal).macOS 13+
accelerateApple Accelerate BLAS (CPU, macOS only).macOS
mklIntel MKL BLAS (CPU, Linux/Windows). Faster on Intel CPUs.Intel oneMKL runtime
aoclAMD AOCL-BLAS / BLIS (CPU). Faster on Zen-class AMD CPUs.AMD AOCL runtime

Multiple CPU backends can coexist: with --features mkl,aocl both will register on startup, the Phase 6b judge profiles each, and the dispatch table picks the winner per (op, dtype, size_class) empirically. The "wrong" backend for a given CPU (MKL on AMD, AOCL on Intel) just loses the profile race — it doesn't break anything, and there's no need to gate via #[cfg(target_arch)] heuristics.

Runtime requirements (where the shared libraries come from)

Cargo features enable the Rust glue. The actual numerical kernels live in vendor-shipped shared libraries that must be resolvable by the OS dynamic loader at runtime — not at compile time. cargo build --features X succeeds without the runtime present; the binary fails on first call when it can't find the DLL / .so / .dylib.

The required runtime library, default install path, and how the OS loader finds it for each feature:

cudanvcuda.dll / libcuda.so from the NVIDIA driver install. The driver installer adds it to system PATH / ld.so config automatically.

cudnncudnn*.dll / libcudnn.so from the CUDA toolkit or a standalone cuDNN install. Add <cuda>/bin to PATH (Windows) or <cuda>/lib64 to LD_LIBRARY_PATH (Linux).

vulkanvulkan-1.dll / libvulkan.so.1 from the GPU driver install. The driver installer adds it to system PATH / ld.so config automatically.

metal / accelerate — built into macOS. Always available.

mklmkl_rt.2.dll / libmkl_rt.so.2. Default install paths: C:\Program Files (x86)\Intel\oneAPI\mkl\<ver> on Windows; /opt/intel/oneapi/mkl/<ver> on Linux. Run setvars.bat / source setvars.sh from the oneAPI install dir, OR add the MKL redist / lib/intel64 directory to PATH / LD_LIBRARY_PATH.

aoclAOCL-LibBlis-Win-dll.dll / libblis.so. Default install paths: C:\Program Files\AMD\AOCL-Windows\amd-blis\lib\LP64 on Windows; /opt/AMD/aocl-linux-*/aocl-blis/lib/LP64 on Linux. Not added to system PATH by the AOCL installer. Add the lib/LP64 directory above to PATH / LD_LIBRARY_PATH manually before running.

If a runtime library is missing or off the loader's search path, you'll see errors like STATUS_DLL_NOT_FOUND (Windows error 0xc0000135) or error while loading shared libraries (Linux). The fix is always "make the directory containing the named library visible to the dynamic loader" — either through PATH/LD_LIBRARY_PATH, the OS config files (/etc/ld.so.conf.d/ on Linux), or by copying the DLL next to your executable on Windows.

Backends self-test on startup. AoclBackend::try_new() runs a 2×2 sgemm to verify the library actually loaded; if the DLL is missing the call returns Err, the backend doesn't register, and the rest of Fuel transparently falls back to other CPU backends. You won't see a hard crash from a missing optional runtime — only an eprintln! from the probe collector.

On Windows, both AoclBackend::try_new and MklBackend::try_new discover the vendor's BLIS / mkl_rt DLL automatically — they look at standard install paths and the AOCL_ROOT / MKLROOT env vars and prepend the matching bin directory to the process's PATH before the load probe. So cargo run --features aocl,onemkl works out of the box on a normal AOCL / oneAPI install without any manual setvars.bat or path-extension shell prep.

Activating empirical backend selection

Compiling with --features aocl,onemkl registers both backends, but by default Tensor::realize_f32() keeps using the portable Rust gemm — exactly as it did before the per-vendor backends existed. To switch on per-op empirical routing, the app calls populate_dispatch_table() once:

use fuel_core::dispatch;

fn main() -> fuel_core::Result<()> {
    // Option 1: blocking on the main thread. First run measures every
    // backend × op × size_class (~10–60s depending on hardware) and
    // persists the profile to disk. Every subsequent run loads from
    // disk in sub-millisecond.
    dispatch::populate_dispatch_table()?;

    // Option 2: background thread. Routing kicks in once the judge
    // returns; the first few realize calls fall through to the
    // portable CPU baseline, which is fine.
    std::thread::spawn(|| {
        let _ = dispatch::populate_dispatch_table();
    });

    // Option 3: skip the call entirely. realize_f32 keeps using the
    // portable CPU path; no behaviour change. The disk-cache lazy-load
    // means a previous process's `populate_dispatch_table()` is still
    // honored — `dispatch::cached()` quietly loads it on first use.

    // ... your model code uses Tensor::realize_f32() as normal ...
    Ok(())
}

Once a dispatch table is cached, every Tensor::realize_f32() call consults it per op. On a Zen-class AMD CPU with both AOCL and oneMKL enabled, this typically picks AOCL or MKL (whichever wins the empirical race that run) for matmul-heavy work and stays on the portable backend for the few percent that's elementwise. No code changes downstream — the realize_f32() call site is identical to the no-routing default.

If a previous profile becomes stale (driver upgrade, BLAS lib swap, OS kernel update with measurably different behaviour), call dispatch::invalidate(). The next populate_dispatch_table() re-runs the judge and overwrites the persisted profile.

Where to download the vendor runtimes

To add a GPU backend when running examples:

# NVIDIA GPU
cargo run --features cuda --example <name> --release

# Apple Silicon
cargo run --features metal --example <name> --release

# CPU only (no GPU toolkit needed)
cargo run --example <name> --release

To add GPU support to your own project:

# Cargo.toml
[dependencies]
fuel-core = { version = "0.10.2", features = ["cuda"] }   # NVIDIA
fuel-core = { version = "0.10.2", features = ["metal"] }  # Apple
fuel-core = { version = "0.10.2" }                        # CPU only

There are also some wasm examples for whisper and llama2.c. You can either build them with trunk or try them online: whisper, llama2, T5, Phi-1.5, and Phi-2, Segment Anything Model.

For LLaMA2, run the following command to retrieve the weight files and start a test server:

cd fuel-wasm-examples/llama2-c
wget https://huggingface.co/spaces/lmz/candle-llama2/resolve/main/model.bin
wget https://huggingface.co/spaces/lmz/candle-llama2/resolve/main/tokenizer.json
trunk serve --release --port 8081

And then head over to http://localhost:8081/.

Useful External Resources

  • candle-tutorial: A very detailed tutorial showing how to convert a PyTorch model to Fuel.
  • candle-lora: Efficient and ergonomic LoRA implementation for Fuel. fuel-lora has
    out-of-the-box LoRA support for many models from Fuel, which can be found here.
  • candle-video: Rust library for text-to-video generation (LTX-Video and related models) built on Candle, focused on fast, Python-free inference.
  • optimisers: A collection of optimisers including SGD with momentum, AdaGrad, AdaDelta, AdaMax, NAdam, RAdam, and RMSprop.
  • candle-vllm: Efficient platform for inference and serving local LLMs including an OpenAI compatible API server.
  • candle-ext: An extension library to Candle that provides PyTorch functions not currently available in Candle.
  • candle-coursera-ml: Implementation of ML algorithms from Coursera's Machine Learning Specialization course.
  • kalosm: A multi-modal meta-framework in Rust for interfacing with local pre-trained models with support for controlled generation, custom samplers, in-memory vector databases, audio transcription, and more.
  • candle-sampling: Sampling techniques for Candle.
  • gpt-from-scratch-rs: A port of Andrej Karpathy's Let's build GPT tutorial on YouTube showcasing the Fuel API on a toy problem.
  • candle-einops: A pure rust implementation of the python einops library.
  • atoma-infer: A Rust library for fast inference at scale, leveraging FlashAttention2 for efficient attention computation, PagedAttention for efficient KV-cache memory management, and multi-GPU support. It is OpenAI api compatible.
  • llms-from-scratch-rs: A comprehensive Rust translation of the code from Sebastian Raschka's Build an LLM from Scratch book.
  • vllm.rs: A minimalist vLLM implementation in Rust based on Candle.

If you have an addition to this list, please submit a pull request.

Features

  • Simple syntax, looks and feels like PyTorch.
  • Backends.
    • Optimized CPU backend with optional MKL support for x86 and Accelerate for macs.
    • CUDA backend for efficiently running on GPUs, multiple GPU distribution via NCCL.
    • WASM support, run your models in a browser.
  • Included models.
    • Language Models.
      • LLaMA v1, v2, and v3 with variants such as SOLAR-10.7B.
      • Falcon.
      • StarCoder, StarCoder2.
      • Phi 1, 1.5, 2, and 3.
      • Mamba, Minimal Mamba
      • Gemma v1 2b and 7b+, v2 2b and 9b.
      • Mistral 7b v0.1.
      • Mixtral 8x7b v0.1.
      • StableLM-3B-4E1T, StableLM-2-1.6B, Stable-Code-3B.
      • Replit-code-v1.5-3B.
      • Bert.
      • Yi-6B and Yi-34B.
      • Qwen1.5, Qwen1.5 MoE, Qwen3 MoE.
      • RWKV v5 and v6.
    • Quantized LLMs.
      • Llama 7b, 13b, 70b, as well as the chat and code variants.
      • Mistral 7b, and 7b instruct.
      • Mixtral 8x7b.
      • Zephyr 7b a and b (Mistral-7b based).
      • OpenChat 3.5 (Mistral-7b based).
      • Qwen3 MoE (16B-A3B, 32B-A3B)
    • Text to text.
      • T5 and its variants: FlanT5, UL2, MADLAD400 (translation), CoEdit (Grammar correction).
      • Marian MT (Machine Translation).
    • Text to image.
      • Stable Diffusion v1.5, v2.1, XL v1.0.
      • Wurstchen v2.
    • Image to text.
      • BLIP.
      • TrOCR.
    • Audio.
      • Whisper, multi-lingual speech-to-text.
      • EnCodec, audio compression model.
      • MetaVoice-1B, text-to-speech model.
      • Parler-TTS, text-to-speech model.
    • Computer Vision Models.
      • DINOv2, ConvMixer, EfficientNet, ResNet, ViT, VGG, RepVGG, ConvNeXT, ConvNeXTv2, MobileOne, EfficientVit (MSRA), MobileNetv4, Hiera, FastViT.
      • yolo-v3, yolo-v8.
      • Segment-Anything Model (SAM).
      • SegFormer.
  • File formats: load models from safetensors, npz, ggml, or PyTorch files.
  • Serverless (on CPU), small and fast deployments.
  • Quantization support using the llama.cpp quantized types.

How to use

Cheatsheet:

Using PyTorchUsing Fuel
Creationtorch.Tensor([[1, 2], [3, 4]])Tensor::new(&[[1f32, 2.], [3., 4.]], &Device::Cpu)?
Creationtorch.zeros((2, 2))Tensor::zeros((2, 2), DType::F32, &Device::Cpu)?
Indexingtensor[:, :4]tensor.i((.., ..4))?
Operationstensor.view((2, 2))tensor.reshape((2, 2))?
Operationsa.matmul(b)a.matmul(&b)?
Arithmetica + b&a + &b
Devicetensor.to(device="cuda")tensor.to_device(&fuel_core::cuda_backend::new_device(0)?)?
Dtypetensor.to(dtype=torch.float16)tensor.to_dtype(&DType::F16)?
Savingtorch.save({"A": A}, "model.bin")fuel::safetensors::save(&HashMap::from([("A", A)]), "model.safetensors")?
Loadingweights = torch.load("model.bin")fuel::safetensors::load("model.safetensors", &device)

Structure

FAQ

Why should I use Fuel?

Fuel's core goal is to make serverless inference possible. Full machine learning frameworks like PyTorch are very large, which makes creating instances on a cluster slow. Fuel allows deployment of lightweight binaries.

Secondly, Fuel lets you remove Python from production workloads. Python overhead can seriously hurt performance, and the GIL is a notorious source of headaches.

Finally, Rust is cool! A lot of the HF ecosystem already has Rust crates, like safetensors and tokenizers.

Other ML frameworks

  • dfdx is a formidable crate, with shapes being included in types. This prevents a lot of headaches by getting the compiler to complain about shape mismatches right off the bat. However, we found that some features still require nightly, and writing code can be a bit daunting for non rust experts.

    We're leveraging and contributing to other core crates for the runtime so hopefully both crates can benefit from each other.

  • burn is a general crate that can leverage multiple backends so you can choose the best engine for your workload.

  • tch-rs Bindings to the torch library in Rust. Extremely versatile, but they bring in the entire torch library into the runtime. The main contributor of tch-rs is also involved in the development of fuel.

Common Errors

Missing symbols when compiling with the mkl feature.

If you get some missing symbols when compiling binaries/tests using the mkl or accelerate features, e.g. for mkl you get:

  = note: /usr/bin/ld: (....o): in function `blas::sgemm':
          .../blas-0.22.0/src/lib.rs:1944: undefined reference to `sgemm_' collect2: error: ld returned 1 exit status

  = note: some `extern` functions couldn't be found; some native libraries may need to be installed or have their path specified
  = note: use the `-l` flag to specify native libraries to link
  = note: use the `cargo:rustc-link-lib` directive to specify the native libraries to link with Cargo

or for accelerate:

Undefined symbols for architecture arm64:
            "_dgemm_", referenced from:
                fuel_core::accelerate::dgemm::h1b71a038552bcabe in libfuel_core...
            "_sgemm_", referenced from:
                fuel_core::accelerate::sgemm::h2cf21c592cba3c47 in libfuel_core...
          ld: symbol(s) not found for architecture arm64

This is likely due to a missing linker flag that was needed to enable the mkl library. You can try adding the following for mkl at the top of your binary:

extern crate intel_mkl_src;

or for accelerate:

extern crate accelerate_src;

Cannot run the LLaMA examples: access to source requires login credentials

Error: request error: https://huggingface.co/meta-llama/Llama-2-7b-hf/resolve/main/tokenizer.json: status code 401

This is likely because you're not permissioned for the LLaMA-v2 model. To fix this, you have to register on the huggingface-hub, accept the LLaMA-v2 model conditions, and set up your authentication token. See issue #350 for more details.

Docker build

When building CUDA kernels inside a Dockerfile, nvidia-smi cannot be used to auto-detect compute capability.

You must explicitly set CUDA_COMPUTE_CAP, for example:

FROM nvidia/cuda:12.9.0-devel-ubuntu22.04

# Install git and curl
RUN set -eux; \
  apt-get update; \
  apt-get install -y curl git ca-certificates;

# Install Rust
RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y

# Clone fuel repo
RUN git clone https://github.com/ciresnave/fuel.git

# Set compute capability for the build
ARG CUDA_COMPUTE_CAP=90
ENV CUDA_COMPUTE_CAP=${CUDA_COMPUTE_CAP}

# Build with explicit compute cap
WORKDIR /app
COPY . .
RUN cargo build --release features cuda

Compiling with flash-attention fails

/usr/include/c++/11/bits/std_function.h:530:146: error: parameter packs not expanded with ‘...’:

This is a bug in gcc-11 triggered by the Cuda compiler. To fix this, install a different, supported gcc version - for example gcc-10, and specify the path to the compiler in the NVCC_CCBIN environment variable.

env NVCC_CCBIN=/usr/lib/gcc/x86_64-linux-gnu/10 cargo ...

Extremely slow model load time with WSL

This may be caused by the models being loaded from /mnt/c, more details on stackoverflow.

Tracking down errors

You can set RUST_BACKTRACE=1 to be provided with backtraces when a fuel error is generated.

CudaRC error

If you encounter an error like this one called Result::unwrap()on anErr value: LoadLibraryExW { source: Os { code: 126, kind: Uncategorized, message: "The specified module could not be found." } } on windows. To fix copy and rename these 3 files (make sure they are in path). The paths depend on your cuda version. c:\Windows\System32\nvcuda.dll -> cuda.dll c:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4\bin\cublas64_12.dll -> cublas.dll c:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.4\bin\curand64_10.dll -> curand.dll

Contributors

(top 30 of 252)

ciresnave

2,635 commits

LaurentMazare

1,603 commits

Narsil

275 commits

ciresnave-bot

166 commits

Languages

Rust

94.4%

Metal

2.4%

Slang

1.5%