High performance, multimodal-native engine for AI workloads.
See the codeA high-performance, multimodal-native engine for AI workloads
Vane Data is a high-performance, multimodal-native data engine for AI workloads. Built on a fork of DuckDB, it extends the core execution engine with native multimodal processing and a unified framework for local and distributed execution.

Vane supports Python 3.10 through 3.14. Python 3.12 is recommended and is the primary development version.
Install the vane-ai package from PyPI:
pip install vane-ai
For more details, see the Installation Guide.
Follow the Quickstart guide to build and run your first Vane pipeline.
Hardware configuration: 1 node, 36 CPU cores, 64 GB memory, and 1× NVIDIA GeForce RTX 2080 Ti (22 GB VRAM).
We use the Ray Data benchmark suite to compare Vane with Ray Data and Daft. The benchmark source code is included in this repository.

The Ray runner targets distributed workloads. The current results are single-node only; validation on the multi-node environments used in the Ray Data benchmarks is still pending.
See the benchmarking page for detailed results.
Contributions and collaborations are welcome. Contribution guidelines and community channels will be published as the project opens further.
Vane is distributed under the Apache License 2.0. See LICENSE and NOTICE for details and third-party attributions.
Vane Data is built on top of DuckDB and inspired by infrastructure systems such as Ray Data, Daft, and Trino.
Special thanks to these projects.
Give Vane a ⭐️ if it helps you!
(top 30 of 36)
Python
85.2%
C++
14.2%
High performance, multimodal-native engine for AI workloads.
See the codeA high-performance, multimodal-native engine for AI workloads
Vane Data is a high-performance, multimodal-native data engine for AI workloads. Built on a fork of DuckDB, it extends the core execution engine with native multimodal processing and a unified framework for local and distributed execution.

Vane supports Python 3.10 through 3.14. Python 3.12 is recommended and is the primary development version.
Install the vane-ai package from PyPI:
pip install vane-ai
For more details, see the Installation Guide.
Follow the Quickstart guide to build and run your first Vane pipeline.
Hardware configuration: 1 node, 36 CPU cores, 64 GB memory, and 1× NVIDIA GeForce RTX 2080 Ti (22 GB VRAM).
We use the Ray Data benchmark suite to compare Vane with Ray Data and Daft. The benchmark source code is included in this repository.

The Ray runner targets distributed workloads. The current results are single-node only; validation on the multi-node environments used in the Ray Data benchmarks is still pending.
See the benchmarking page for detailed results.
Contributions and collaborations are welcome. Contribution guidelines and community channels will be published as the project opens further.
Vane is distributed under the Apache License 2.0. See LICENSE and NOTICE for details and third-party attributions.
Vane Data is built on top of DuckDB and inspired by infrastructure systems such as Ray Data, Daft, and Trino.
Special thanks to these projects.
Give Vane a ⭐️ if it helps you!
(top 30 of 36)
Python
85.2%
C++
14.2%