Vx: one language, every chip.
See the codeOne language, every chip.
A heterogeneous-systems programming language that puts hardware topology, memory placement, and reachability directly in the type system.
Why Vx • 30-Second Demo • Quickstart • Core Pillars • Architecture • Status & Limitations • Documentation • Discord
Modern high-performance programs are rarely confined to a single CPU. They orchestrate work across host DRAM, PCIe buses, GPU high-bandwidth memory (HBM), and specialized accelerators like NPUs or the Apple Neural Engine (ANE).
In conventional languages (CUDA, C++, Python), hardware topology and memory spaces are invisible to the compiler. If host code reads a device pointer, or if an asynchronous transfer is not awaited, your program fails at runtime: with a silent corruption, a fatal segmentation fault, or an out-of-memory (OOM) crash hours into a distributed training job.
Vx eliminates this class of bugs at compile time.
Where a value lives (Memory::CPU_DRAM, Memory::GPU_HBM, Memory::NPU_HBM) and where code executes (Topology::GPU[0], Topology::NPU[0], Topology::ANE) are first-class types. The compiler checks memory capacity, bus bandwidth, and address visibility before code ever touches silicon:
| Challenge | Today (CUDA / C++ / PyTorch) | The Vx Way |
|---|---|---|
| Invalid device memory access | Host reads GPU pointer $\rightarrow$ runtime segfault (cudaErrorIllegalAddress) | Compile Error E6003: Compiler proves host cannot address device space and suggests the missing transfer |
| Memory capacity exhaustion | Runtime OOM crash when a tensor working set exceeds VRAM | Compile Error E6009 / E6010: Compiler checks memory capacity ahead of time against declared hardware limits |
| Targeting multi-accelerator nodes | Fragmented across CUDA, Metal, OpenCL, and proprietary vendor runtimes | Unified syntax: spawn on(Topology::...) with declarative hardware definitions in fleet/ |
| Asynchronous data hazards | Race conditions and manual stream synchronization bugs | Formally verified seam contracts: Transfer bounds and visibility are proven by a Z3 solver |
Here is a complete matrix multiplication that prepares data on the host, stages it into an accelerator's HBM, executes on the device, and brings the result back:
fn main() -> i32 {
// 1. Allocate and initialize host matrices
let a = Tensor<f32, [4, 4]>::fill(2.0);
let b = Tensor<f32, [4, 4]>::fill(3.0);
let mut c = Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();
// 2. Explicitly stage data across the interconnect into accelerator HBM
let a_npu = transfer(a, Memory::NPU_HBM);
let b_npu = transfer(b, Memory::NPU_HBM);
// 3. Dispatch execution onto the declared topology
spawn on(Topology::NPU[0]) {
c = a_npu @ b_npu;
}
// 4. Retrieve result back to host DRAM
let c_host = transfer(c, Memory::CPU_DRAM);
print(c_host[0][0]); // 24
return 0;
}
Run it with the built-in JIT:
$ vxc matmul.vx
24
If you forget to transfer the input a and attempt to read it inside the spawn on(Topology::NPU[0]) block, the compiler refuses the program immediately:
Error[E6003] at matmul.vx:12:9: 'a' lives in CPU_DRAM but NPU[0] sees only [NPU_HBM];
insert an explicit transfer to NPU_HBM (cost 50 on the declared path)
The compiler proves that NPU[0] cannot address CPU_DRAM, finds the cheapest legal route across the declared bus edges, calculates the transfer cost, and provides the fix.
brew install llvm z3 cmakeapt install z3 lld cmake# Clone the repository
git clone https://github.com/vx-lang/Vx.git
cd Vx
# Locate LLVM and write local configuration
./setup.sh
# Source environment variables (required before running cargo)
source config.local
# Build the release compiler
cargo build --release
# Run a smoke test with the JIT compiler (default action)
./target/release/vxc examples/docker_smoke.vx
# Compile and check against a declared NVIDIA H100 machine (works on any host!)
./target/release/vxc --host default --machine fleet/h100-sxm.vx examples/docker_smoke.vx -o smoke_h100
# Inspect capacity admission and routing costs as JSON
./target/release/vxc --host default --machine fleet/h100-sxm.vx examples/docker_smoke.vx --diagnostics-json out.json
# Inspect the generated MLIR
./target/release/vxc examples/docker_smoke.vx --action emit-mlir
source config.local
cargo test
Over 530 unit tests and 40 integration suites verify placement rules, differential CUDA checks, and determinism.
A tensor's type in Vx carries its element type, dimensions, and its residency domain: Tensor<f32, [4, 4], Memory::GPU_HBM>.
The compiler enforces strict invariants before codegen:
| Check | Diagnostic | Rules Out |
|---|---|---|
| Address-space visibility | E6003 | Host code reading device memory, or an accelerator accessing inaccessible address spaces |
| Capacity admission | E6009, E6010 | Single tensors or multi-tensor working sets that exceed available memory |
| Transfer reachability | E6002 | Copying between memory domains with no declared hardware edge |
| Optimal transfer routing | Cost Model | Sub-optimal routes; automatically prices containment hops across memory hierarchies |
| Seam contracts | Prover (z3) | Reading an asynchronous transfer buffer before it has been synchronized |
| Linear buffers | Borrow Checker | Use-after-move of consumed device memory buffers |
A differential test suite pairs each check against CUDA on an NVIDIA A100: errors that CUDA detects as runtime aborts, failed cudaMalloc allocations, or segfaults are caught by Vx at compile time.
fleet/)Instead of hardcoding memory sizes and bus links, machines are declared as clean specification files:
Memory HBM { capacity: 80 GiB, bandwidth: 3.35 TB/s, managed: explicit, scope: device }
Memory L2 { within: Memory::HBM, capacity: 50 MiB, bandwidth: 12 TB/s, managed: cached }
Memory SMEM { within: Memory::L2, capacity: 228 KiB, bandwidth: 128 B/cyc,
clock: 1.98 GHz, replicas: 132, granule: 1 KiB, scope: sm }
Topology Device {
arch: nvptx64,
memory: Memory::HBM,
transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
}
The fleet/ directory includes 12 validated specifications:
Because vxc cross-compiles against declared machines (--machine fleet/<sku>.vx), you can verify and compile binaries for an 8-GPU H100 cluster directly from an M4 MacBook Air without renting cloud GPUs.
Traditional compiler frontends frequently bottleneck on single-threaded symbol resolution or lock contention (Mutex/RwLock). Vx introduces an 8-phase parallel architecture:
[u64; 4] content hash. Symbol lookup involves zero pointer chasing or string hashing.GlobalSession) guarantees that worker threads never contend.For a deep dive into the design, see the Parallel Compiler Architecture.
Vx includes native tensor abstractions with rank and extent tracking, dynamic dimensions (Tensor<f32, [?, ?]>), matrix multiplication (@), and built-in automatic differentiation (grad, vjp, jvp) lowering through an optional Enzyme MLIR plugin:
fn loss(x: f32) -> f32 {
return x * x;
}
fn main() -> i32 {
let dx: f32 = grad(loss, 3.0); // 6.0
print(dx);
return 0;
}
For a full real-world demonstration, explore examples/llama.vx, which implements a complete, clean Llama 2 forward pass with tensor abstractions.
flowchart TD
SRC["Source Program (*.vx)"] --> PARSER["Parallel Frontend\n(256-bit GID content hashes)"]
FLEET["Machine Declaration (fleet/*.vx)"] --> CHECKER["Topology & Memory Algebra Engine\n(Capacity E6009, Visibility E6003, Z3 Seams)"]
PARSER --> CHECKER
CHECKER --> HIR["Flat HIR Instruction Stream"]
HIR --> MLIR["MLIR (Custom vx Dialect)"]
MLIR --> CPU["Host CPU (LLVM IR -> JIT or Object File)"]
MLIR --> GPU["NVIDIA GPU (NVVM -> PTX -> cuBLAS)"]
MLIR --> ANE["Apple Silicon (CoreML Plugin -> Neural Engine)"]
MLIR --> REMOTE["Remote Nodes (Wire Protocol over SSH / TCP)"]
| Target | Lowering Pipeline | Tested & Verified Workloads |
|---|---|---|
| CPU (x86_64, ARM64) | MLIR $\rightarrow$ LLVM IR $\rightarrow$ Native (JIT or .o) | Full test suite and standard library |
| NVIDIA GPU | MLIR $\rightarrow$ NVVM $\rightarrow$ PTX payload via CUDA driver | Disaggregated prefill/decode split, cuBLAS GEMMs on A100 & H100 |
| Apple Silicon (ANE) | CoreML dispatch from native plugin | FP16 512×512 matmul on Neural Engine, FP32 dispatched to CPU |
| Remote Node | Wire protocol carrying memref descriptors | ARM64 laptop orchestrating an x86_64 remote worker over SSH |
Release Version: Vx is currently in v0.0.2. The syntax and core type-system checks are stable. The placement, routing, and memory algebra systems are verified by active test suites. However, as an early research systems compiler, several language features are actively being built.
| Category | Status in v0.0.2 | Tracking Issue / Reference |
|---|---|---|
| Memory Algebra & Topology | ✅ Implemented, tested, diagnostic codes active | docs/memory_algebra.md |
| JIT & AOT Cross-Compilation | ✅ Fully supported via LLVM and declared machines | ROADMAP.md |
| Control Flow | ⚠️ for loops over ranges and recursion work; while loops in development | #506 |
Overlapped spawn on | ⚠️ spawn on(...) { ... } is sequential by design; the non-blocking dispatch that overlaps it with host work is still to be built | docs/spawn_on.md |
| Standard Library | ⚠️ Core modules (io, math, vec, fs, net, time) working; package manager in progress | #487 |
| Automatic Memory Cleanup | ⚠️ Linear types and explicit free() available; RAII destructors in development | #495 |
Autodiff (grad) | ⚠️ Enzyme MLIR integration functional; discrete function checks in progress | #503 |
Full tracking is available on our GitHub Issue Tracker.
vx-analyzer: Language Server Protocol (LSP) providing diagnostics, hover tooltips, and go-to-definition.vscode-vx: Official VS Code extension providing syntax highlighting and LSP integration.vx-format: Official AST-aware code formatter for .vx files.vx-opt: Specialized driver for testing and inspecting passes on the custom vx MLIR dialect.src/ Compiler implementation in Rust
lexer, parser/ Source parsing and interface extraction
syntax/ AST, type system, and topology declarations
hir/ Parallel type checking, borrow checking, memory algebra
codegen/ MLIR generation
dialect/ Custom vx MLIR dialect and lowering passes (C++)
gid.rs 256-bit content-hashed Global Identifiers
pipeline.rs Zero-lock 8-phase parallel frontend
runtime/ Dispatch runtimes (Host CPU, NVIDIA CUDA, Apple CoreML, Remote Worker)
fleet/ Declarative machine descriptions (A100, H100, B200, MI300X, M4, etc.)
stdlib/ Standard library modules (tensor, simd, vec, io, math, fs, etc.)
examples/ Example programs (including Llama 2 forward pass in llama.vx)
docs/ Language specification, tutorials, and architecture designs
vx-analyzer/ Language server (LSP)
vscode-vx/ VS Code extension
tests/ Unit, integration, middle-end, and differential CUDA test suites
If you use Vx in your research or systems work, please cite:
@misc{vx2026,
author = {Aditya Kumar},
title = {{Vx}: a systems programming language for heterogeneous computing},
year = {2026},
howpublished = {\url{https://vxlang.org}},
note = {Version 0.0.2. Source at \url{https://github.com/vx-lang/Vx}}
}
Vx is distributed under the Apache License 2.0 with LLVM Exceptions. See LICENSE for details.
116 followers · starred Sep 2026
96 followers · starred Sep 2026
34 followers · starred Sep 2026
73 followers · starred Sep 2026
Rust
74.7%
C++
8.9%
HTML
4.5%
Shell
4.1%
Python
3.9%
Cuda
1.3%
Objective-C++
1.1%
Vx: one language, every chip.
See the codeOne language, every chip.
A heterogeneous-systems programming language that puts hardware topology, memory placement, and reachability directly in the type system.
Why Vx • 30-Second Demo • Quickstart • Core Pillars • Architecture • Status & Limitations • Documentation • Discord
Modern high-performance programs are rarely confined to a single CPU. They orchestrate work across host DRAM, PCIe buses, GPU high-bandwidth memory (HBM), and specialized accelerators like NPUs or the Apple Neural Engine (ANE).
In conventional languages (CUDA, C++, Python), hardware topology and memory spaces are invisible to the compiler. If host code reads a device pointer, or if an asynchronous transfer is not awaited, your program fails at runtime: with a silent corruption, a fatal segmentation fault, or an out-of-memory (OOM) crash hours into a distributed training job.
Vx eliminates this class of bugs at compile time.
Where a value lives (Memory::CPU_DRAM, Memory::GPU_HBM, Memory::NPU_HBM) and where code executes (Topology::GPU[0], Topology::NPU[0], Topology::ANE) are first-class types. The compiler checks memory capacity, bus bandwidth, and address visibility before code ever touches silicon:
| Challenge | Today (CUDA / C++ / PyTorch) | The Vx Way |
|---|---|---|
| Invalid device memory access | Host reads GPU pointer $\rightarrow$ runtime segfault (cudaErrorIllegalAddress) | Compile Error E6003: Compiler proves host cannot address device space and suggests the missing transfer |
| Memory capacity exhaustion | Runtime OOM crash when a tensor working set exceeds VRAM | Compile Error E6009 / E6010: Compiler checks memory capacity ahead of time against declared hardware limits |
| Targeting multi-accelerator nodes | Fragmented across CUDA, Metal, OpenCL, and proprietary vendor runtimes | Unified syntax: spawn on(Topology::...) with declarative hardware definitions in fleet/ |
| Asynchronous data hazards | Race conditions and manual stream synchronization bugs | Formally verified seam contracts: Transfer bounds and visibility are proven by a Z3 solver |
Here is a complete matrix multiplication that prepares data on the host, stages it into an accelerator's HBM, executes on the device, and brings the result back:
fn main() -> i32 {
// 1. Allocate and initialize host matrices
let a = Tensor<f32, [4, 4]>::fill(2.0);
let b = Tensor<f32, [4, 4]>::fill(3.0);
let mut c = Tensor<f32, [4, 4], Memory::NPU_HBM>::uninit();
// 2. Explicitly stage data across the interconnect into accelerator HBM
let a_npu = transfer(a, Memory::NPU_HBM);
let b_npu = transfer(b, Memory::NPU_HBM);
// 3. Dispatch execution onto the declared topology
spawn on(Topology::NPU[0]) {
c = a_npu @ b_npu;
}
// 4. Retrieve result back to host DRAM
let c_host = transfer(c, Memory::CPU_DRAM);
print(c_host[0][0]); // 24
return 0;
}
Run it with the built-in JIT:
$ vxc matmul.vx
24
If you forget to transfer the input a and attempt to read it inside the spawn on(Topology::NPU[0]) block, the compiler refuses the program immediately:
Error[E6003] at matmul.vx:12:9: 'a' lives in CPU_DRAM but NPU[0] sees only [NPU_HBM];
insert an explicit transfer to NPU_HBM (cost 50 on the declared path)
The compiler proves that NPU[0] cannot address CPU_DRAM, finds the cheapest legal route across the declared bus edges, calculates the transfer cost, and provides the fix.
brew install llvm z3 cmakeapt install z3 lld cmake# Clone the repository
git clone https://github.com/vx-lang/Vx.git
cd Vx
# Locate LLVM and write local configuration
./setup.sh
# Source environment variables (required before running cargo)
source config.local
# Build the release compiler
cargo build --release
# Run a smoke test with the JIT compiler (default action)
./target/release/vxc examples/docker_smoke.vx
# Compile and check against a declared NVIDIA H100 machine (works on any host!)
./target/release/vxc --host default --machine fleet/h100-sxm.vx examples/docker_smoke.vx -o smoke_h100
# Inspect capacity admission and routing costs as JSON
./target/release/vxc --host default --machine fleet/h100-sxm.vx examples/docker_smoke.vx --diagnostics-json out.json
# Inspect the generated MLIR
./target/release/vxc examples/docker_smoke.vx --action emit-mlir
source config.local
cargo test
Over 530 unit tests and 40 integration suites verify placement rules, differential CUDA checks, and determinism.
A tensor's type in Vx carries its element type, dimensions, and its residency domain: Tensor<f32, [4, 4], Memory::GPU_HBM>.
The compiler enforces strict invariants before codegen:
| Check | Diagnostic | Rules Out |
|---|---|---|
| Address-space visibility | E6003 | Host code reading device memory, or an accelerator accessing inaccessible address spaces |
| Capacity admission | E6009, E6010 | Single tensors or multi-tensor working sets that exceed available memory |
| Transfer reachability | E6002 | Copying between memory domains with no declared hardware edge |
| Optimal transfer routing | Cost Model | Sub-optimal routes; automatically prices containment hops across memory hierarchies |
| Seam contracts | Prover (z3) | Reading an asynchronous transfer buffer before it has been synchronized |
| Linear buffers | Borrow Checker | Use-after-move of consumed device memory buffers |
A differential test suite pairs each check against CUDA on an NVIDIA A100: errors that CUDA detects as runtime aborts, failed cudaMalloc allocations, or segfaults are caught by Vx at compile time.
fleet/)Instead of hardcoding memory sizes and bus links, machines are declared as clean specification files:
Memory HBM { capacity: 80 GiB, bandwidth: 3.35 TB/s, managed: explicit, scope: device }
Memory L2 { within: Memory::HBM, capacity: 50 MiB, bandwidth: 12 TB/s, managed: cached }
Memory SMEM { within: Memory::L2, capacity: 228 KiB, bandwidth: 128 B/cyc,
clock: 1.98 GHz, replicas: 132, granule: 1 KiB, scope: sm }
Topology Device {
arch: nvptx64,
memory: Memory::HBM,
transfer Memory::CPU_DRAM -> Memory::HBM : 63 GB/s,
}
The fleet/ directory includes 12 validated specifications:
Because vxc cross-compiles against declared machines (--machine fleet/<sku>.vx), you can verify and compile binaries for an 8-GPU H100 cluster directly from an M4 MacBook Air without renting cloud GPUs.
Traditional compiler frontends frequently bottleneck on single-threaded symbol resolution or lock contention (Mutex/RwLock). Vx introduces an 8-phase parallel architecture:
[u64; 4] content hash. Symbol lookup involves zero pointer chasing or string hashing.GlobalSession) guarantees that worker threads never contend.For a deep dive into the design, see the Parallel Compiler Architecture.
Vx includes native tensor abstractions with rank and extent tracking, dynamic dimensions (Tensor<f32, [?, ?]>), matrix multiplication (@), and built-in automatic differentiation (grad, vjp, jvp) lowering through an optional Enzyme MLIR plugin:
fn loss(x: f32) -> f32 {
return x * x;
}
fn main() -> i32 {
let dx: f32 = grad(loss, 3.0); // 6.0
print(dx);
return 0;
}
For a full real-world demonstration, explore examples/llama.vx, which implements a complete, clean Llama 2 forward pass with tensor abstractions.
flowchart TD
SRC["Source Program (*.vx)"] --> PARSER["Parallel Frontend\n(256-bit GID content hashes)"]
FLEET["Machine Declaration (fleet/*.vx)"] --> CHECKER["Topology & Memory Algebra Engine\n(Capacity E6009, Visibility E6003, Z3 Seams)"]
PARSER --> CHECKER
CHECKER --> HIR["Flat HIR Instruction Stream"]
HIR --> MLIR["MLIR (Custom vx Dialect)"]
MLIR --> CPU["Host CPU (LLVM IR -> JIT or Object File)"]
MLIR --> GPU["NVIDIA GPU (NVVM -> PTX -> cuBLAS)"]
MLIR --> ANE["Apple Silicon (CoreML Plugin -> Neural Engine)"]
MLIR --> REMOTE["Remote Nodes (Wire Protocol over SSH / TCP)"]
| Target | Lowering Pipeline | Tested & Verified Workloads |
|---|---|---|
| CPU (x86_64, ARM64) | MLIR $\rightarrow$ LLVM IR $\rightarrow$ Native (JIT or .o) | Full test suite and standard library |
| NVIDIA GPU | MLIR $\rightarrow$ NVVM $\rightarrow$ PTX payload via CUDA driver | Disaggregated prefill/decode split, cuBLAS GEMMs on A100 & H100 |
| Apple Silicon (ANE) | CoreML dispatch from native plugin | FP16 512×512 matmul on Neural Engine, FP32 dispatched to CPU |
| Remote Node | Wire protocol carrying memref descriptors | ARM64 laptop orchestrating an x86_64 remote worker over SSH |
Release Version: Vx is currently in v0.0.2. The syntax and core type-system checks are stable. The placement, routing, and memory algebra systems are verified by active test suites. However, as an early research systems compiler, several language features are actively being built.
| Category | Status in v0.0.2 | Tracking Issue / Reference |
|---|---|---|
| Memory Algebra & Topology | ✅ Implemented, tested, diagnostic codes active | docs/memory_algebra.md |
| JIT & AOT Cross-Compilation | ✅ Fully supported via LLVM and declared machines | ROADMAP.md |
| Control Flow | ⚠️ for loops over ranges and recursion work; while loops in development | #506 |
Overlapped spawn on | ⚠️ spawn on(...) { ... } is sequential by design; the non-blocking dispatch that overlaps it with host work is still to be built | docs/spawn_on.md |
| Standard Library | ⚠️ Core modules (io, math, vec, fs, net, time) working; package manager in progress | #487 |
| Automatic Memory Cleanup | ⚠️ Linear types and explicit free() available; RAII destructors in development | #495 |
Autodiff (grad) | ⚠️ Enzyme MLIR integration functional; discrete function checks in progress | #503 |
Full tracking is available on our GitHub Issue Tracker.
vx-analyzer: Language Server Protocol (LSP) providing diagnostics, hover tooltips, and go-to-definition.vscode-vx: Official VS Code extension providing syntax highlighting and LSP integration.vx-format: Official AST-aware code formatter for .vx files.vx-opt: Specialized driver for testing and inspecting passes on the custom vx MLIR dialect.src/ Compiler implementation in Rust
lexer, parser/ Source parsing and interface extraction
syntax/ AST, type system, and topology declarations
hir/ Parallel type checking, borrow checking, memory algebra
codegen/ MLIR generation
dialect/ Custom vx MLIR dialect and lowering passes (C++)
gid.rs 256-bit content-hashed Global Identifiers
pipeline.rs Zero-lock 8-phase parallel frontend
runtime/ Dispatch runtimes (Host CPU, NVIDIA CUDA, Apple CoreML, Remote Worker)
fleet/ Declarative machine descriptions (A100, H100, B200, MI300X, M4, etc.)
stdlib/ Standard library modules (tensor, simd, vec, io, math, fs, etc.)
examples/ Example programs (including Llama 2 forward pass in llama.vx)
docs/ Language specification, tutorials, and architecture designs
vx-analyzer/ Language server (LSP)
vscode-vx/ VS Code extension
tests/ Unit, integration, middle-end, and differential CUDA test suites
If you use Vx in your research or systems work, please cite:
@misc{vx2026,
author = {Aditya Kumar},
title = {{Vx}: a systems programming language for heterogeneous computing},
year = {2026},
howpublished = {\url{https://vxlang.org}},
note = {Version 0.0.2. Source at \url{https://github.com/vx-lang/Vx}}
}
Vx is distributed under the Apache License 2.0 with LLVM Exceptions. See LICENSE for details.
116 followers · starred Sep 2026
96 followers · starred Sep 2026
34 followers · starred Sep 2026
73 followers · starred Sep 2026
Rust
74.7%
C++
8.9%
HTML
4.5%
Shell
4.1%
Python
3.9%
Cuda
1.3%
Objective-C++
1.1%