An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card.
Verilog
34
1,356 commits
updated Oct 2, 2026
An open-source AI accelerator, developed by AI.
openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?

otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card's
utilization and DRAM bandwidth (right).
openTPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.
The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.
| Model | Weights | Decode, device | Decode, wall | Prefill, device | DRAM while decoding |
|---|---|---|---|---|---|
| LFM2.5-230M | int8 | 59.0 tok/s | 52.3 tok/s | 295.6 tok/s | 14.5 GB/s (85% of peak) |
| LFM2.5-230M | 4-bit, int8 head | 85.8 tok/s | 82.1 tok/s | 335.4 tok/s | 14.1 GB/s (82%) |
| Qwen3-0.6B | int8 | 21.6 tok/s | 21.3 tok/s | 92.1 tok/s | 14.4 GB/s (84%) |
| Qwen3-0.6B | 4-bit, int8 head | 31.3 tok/s | 30.7 tok/s | 103.4 tok/s | 13.9 GB/s (82%) |
| Qwen3.5-0.8B | int8 | 17.6 tok/s | 16.3 tok/s | 61.4 tok/s | 14.5 GB/s (85%) |
| Qwen3.5-0.8B | 4-bit, int8 head | 24.5 tok/s | 23.3 tok/s | 66.7 tok/s | 14.1 GB/s (83%) |
| Gemma 4 E2B | 4-bit, int8 head | 10.57 tok/s | 10.53 tok/s | 32.1 tok/s | 15.6 GB/s (92%) |
| Gemma 4 E2B | 4-bit, 4-bit head | 12.14 tok/s | 12.09 tok/s | 29.9 tok/s | 15.5 GB/s (91%) |
| LFM2-2.6B | int8 | 6.05 tok/s | 6.03 tok/s | 21.4 tok/s | 16.1 GB/s (94%) |
| LFM2-2.6B | 4-bit, int8 head | 10.96 tok/s | 10.93 tok/s | 20.6 tok/s | 15.8 GB/s (93%) |
| SmolLM3-3B | int8 | 5.00 tok/s | 4.99 tok/s | 21.1 tok/s | 16.0 GB/s (94%) |
| SmolLM3-3B | 4-bit, int8 head | 8.74 tok/s | 8.72 tok/s | 22.8 tok/s | 15.7 GB/s (92%) |
| Phi-4-mini (3.8B) | int8 | 3.99 tok/s | 3.98 tok/s | 13.8 tok/s | 16.0 GB/s (94%) |
| Phi-4-mini (3.8B) | 4-bit, int8 head | 6.56 tok/s | 6.55 tok/s | 15.0 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-2B | int8 | 8.02 tok/s | 8.00 tok/s | 38.2 tok/s | 16.0 GB/s (94%) |
| Qwen3.5-2B | 4-bit, int8 head | 12.09 tok/s | 12.03 tok/s | 41.7 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-4B | 4-bit, int8 head | 5.88 tok/s | 5.87 tok/s | 12.9 tok/s | 15.7 GB/s (92%) |
| Gemma 4 E4B | int8, 4-bit head and down 0-23 | 3.78 tok/s | 3.75 tok/s | 14.8 tok/s | 16.0 GB/s (94%) |
Measured on the card: the first three models on 2026-09-29 with the production image
deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and
4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then.
Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%),
at 91-94% of the DRAM peak instead of 82-87%. Qwen3.5-4B's int8 image is over 4 GiB.
tools/qual/perf.py: decode is 64 greedy tokens after a 512-token prompt, with the
host's argmax in the loop (not streamed). "Device" counts only the cycles the accelerator runs;
"wall" adds the host. Prefill is the 512-token prompt, on the device.With the logits streamed back while the card runs (tools/decode_profile.py, 96 tokens), 4-bit
decode is faster, in device / wall tok/s:
The previous production image, se-cand3, was built with the Xilinx MIG, a two-column matrix unit and a 120.755 MHz clock. Measured the same way, the new image:
The earlier images and their numbers are in docs/board.md, section 5.
Mixture-of-experts models bigger than the card's 4 GiB run with their experts streamed from host storage (docs/offload.md, section 10). The card routes each token and computes every expert, and it keeps the experts in per-layer slots in its DRAM. The host only copies missing experts from a pool file into those slots, at the link's rate (section 10.1). Measured on 2026-10-01 with build B (79c5707a), the card's own decode loop picking every token, 4-bit experts, int8 head:
4-bit weights (docs/quant.md) use FP4 values with two-level block scales, 4.25 bits per weight, and keep the LM head in int8 for accuracy. They cut the bytes per token by about a third and raise decode speed by 40% (Qwen3.5) to 45% (Qwen3, LFM2), at a measurable cost in perplexity that docs/quant.md reports per model.
The host is nearly out of the way. For LFM2 and Qwen3 the card runs one decode program compiled once, which reads the position from a register and looks up its own embedding and RoPE rows, and the logits stream back while the card is still running: the host adds 0.17 to 0.30 ms per token on omarchy (0.45 to 1.3 ms on opentpu). Qwen3.5 runs the same way for decode; its prefill still compiles each chunk's program on the host, ahead of the card.
Kernels in ol mlp, attention, full model layers
| @ol.jit
Language + compiler layouts, affine loop addressing, fusion
|
ISA 8 x 32-bit words per instruction
|
ISA simulator <======> RTL same bits, checked by the tests
(Python) (SystemVerilog)
| Vivado bitstream
FPGA card Kintex-7 xc7k480t
| PCIe
Host otpu-chat, otpu-smi, otpu-lens
The machine is deliberately simple. A sequencer issues one instruction per cycle to a few units: DMA moves data, the matrix unit multiplies int8 weights streamed from DRAM, the vector unit does fp32 math, and a quantizer turns results back into int8. There is no cache and no hidden scheduling: every data movement is an instruction, so a trace shows exactly where the cycles go. docs/isa.md describes the whole instruction set.
A kernel looks like this:
from opentpu import language as ol
@ol.jit
def mlp(h, gamma, w_gate, w_up, w_down, out, eps): # simplified; see kernels/mlp.py
x = ol.load(h)
xs = ol.quantize(rmsnorm(x, ol.load(gamma), eps))
g = ol.dot(xs, w_gate)
u = ol.dot(xs, w_up)
a = ol.all_gather(silu(g) * u)
y = ol.all_gather(ol.dot(a, w_down))
if ol.program_id() == 0:
ol.store(out, x + y)
Because every data movement is an instruction, a trace of a run explains its speed. Lens, the profiler, records a run from the RTL, the simulator or the card and opens it in the browser, with a roofline, a timeline and per-instruction tables (docs/lens.md).

Lens replaying part of a Qwen3 decode step. Colours show what each unit is doing in each cycle: busy, waiting on DRAM, or waiting on another instruction.
Everything except the card runs on a laptop.
pip install -e .
pip install pytest torch transformers
python3 -m pytest -q # RTL tests also need Verilator 5
hf download LiquidAI/LFM2.5-230M --local-dir models/LFM2.5-230M
otpu-chat --model lfm2 --backend isa # chat on the simulator
With a card, build the bitstream (make bit in boards/ypcb-00338), load
it over JTAG, then run sudo otpu-setup and otpu-chat --backend board.
docs/board.md walks through the bring-up.
| Command | What it does |
|---|---|
otpu-chat | chat with Qwen3-0.6B, LFM2.5-230M (--model lfm2), Qwen3.5-0.8B (--model qwen35), LFM2-2.6B (lfm2-2.6b), SmolLM3-3B (smollm3), Phi-4-mini (phi4-mini) or Qwen3.5-2B / 4B (qwen35-2b, qwen35-4b) |
otpu-smi | temperature, power, DRAM bandwidth and per-unit utilization |
otpu-lens | record a run and open it in the profiler |
otpu-selftest, otpu-diag | check that the card works |
opentpu/kernels and docs/compiler.md: how a kernel
becomes instructions.opentpu/isasim.py: the simulator, which is the spec.rtl/: the hardware, starting from rtl/top/otpu_top.sv.Issues and pull requests are welcome, and most of the work needs only Python and Verilator, not
an FPGA. Changes to the ISA, the simulator or the RTL must keep python3 -m pytest -q passing,
and performance claims should say how they were measured.
Apache License 2.0. See LICENSE.
An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card.
Verilog
34
1,356 commits
updated Oct 2, 2026
An open-source AI accelerator, developed by AI.
openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?

otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card's
utilization and DRAM bandwidth (right).
openTPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.
The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.
| Model | Weights | Decode, device | Decode, wall | Prefill, device | DRAM while decoding |
|---|---|---|---|---|---|
| LFM2.5-230M | int8 | 59.0 tok/s | 52.3 tok/s | 295.6 tok/s | 14.5 GB/s (85% of peak) |
| LFM2.5-230M | 4-bit, int8 head | 85.8 tok/s | 82.1 tok/s | 335.4 tok/s | 14.1 GB/s (82%) |
| Qwen3-0.6B | int8 | 21.6 tok/s | 21.3 tok/s | 92.1 tok/s | 14.4 GB/s (84%) |
| Qwen3-0.6B | 4-bit, int8 head | 31.3 tok/s | 30.7 tok/s | 103.4 tok/s | 13.9 GB/s (82%) |
| Qwen3.5-0.8B | int8 | 17.6 tok/s | 16.3 tok/s | 61.4 tok/s | 14.5 GB/s (85%) |
| Qwen3.5-0.8B | 4-bit, int8 head | 24.5 tok/s | 23.3 tok/s | 66.7 tok/s | 14.1 GB/s (83%) |
| Gemma 4 E2B | 4-bit, int8 head | 10.57 tok/s | 10.53 tok/s | 32.1 tok/s | 15.6 GB/s (92%) |
| Gemma 4 E2B | 4-bit, 4-bit head | 12.14 tok/s | 12.09 tok/s | 29.9 tok/s | 15.5 GB/s (91%) |
| LFM2-2.6B | int8 | 6.05 tok/s | 6.03 tok/s | 21.4 tok/s | 16.1 GB/s (94%) |
| LFM2-2.6B | 4-bit, int8 head | 10.96 tok/s | 10.93 tok/s | 20.6 tok/s | 15.8 GB/s (93%) |
| SmolLM3-3B | int8 | 5.00 tok/s | 4.99 tok/s | 21.1 tok/s | 16.0 GB/s (94%) |
| SmolLM3-3B | 4-bit, int8 head | 8.74 tok/s | 8.72 tok/s | 22.8 tok/s | 15.7 GB/s (92%) |
| Phi-4-mini (3.8B) | int8 | 3.99 tok/s | 3.98 tok/s | 13.8 tok/s | 16.0 GB/s (94%) |
| Phi-4-mini (3.8B) | 4-bit, int8 head | 6.56 tok/s | 6.55 tok/s | 15.0 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-2B | int8 | 8.02 tok/s | 8.00 tok/s | 38.2 tok/s | 16.0 GB/s (94%) |
| Qwen3.5-2B | 4-bit, int8 head | 12.09 tok/s | 12.03 tok/s | 41.7 tok/s | 15.8 GB/s (92%) |
| Qwen3.5-4B | 4-bit, int8 head | 5.88 tok/s | 5.87 tok/s | 12.9 tok/s | 15.7 GB/s (92%) |
| Gemma 4 E4B | int8, 4-bit head and down 0-23 | 3.78 tok/s | 3.75 tok/s | 14.8 tok/s | 16.0 GB/s (94%) |
Measured on the card: the first three models on 2026-09-29 with the production image
deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and
4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then.
Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%),
at 91-94% of the DRAM peak instead of 82-87%. Qwen3.5-4B's int8 image is over 4 GiB.
tools/qual/perf.py: decode is 64 greedy tokens after a 512-token prompt, with the
host's argmax in the loop (not streamed). "Device" counts only the cycles the accelerator runs;
"wall" adds the host. Prefill is the 512-token prompt, on the device.With the logits streamed back while the card runs (tools/decode_profile.py, 96 tokens), 4-bit
decode is faster, in device / wall tok/s:
The previous production image, se-cand3, was built with the Xilinx MIG, a two-column matrix unit and a 120.755 MHz clock. Measured the same way, the new image:
The earlier images and their numbers are in docs/board.md, section 5.
Mixture-of-experts models bigger than the card's 4 GiB run with their experts streamed from host storage (docs/offload.md, section 10). The card routes each token and computes every expert, and it keeps the experts in per-layer slots in its DRAM. The host only copies missing experts from a pool file into those slots, at the link's rate (section 10.1). Measured on 2026-10-01 with build B (79c5707a), the card's own decode loop picking every token, 4-bit experts, int8 head:
4-bit weights (docs/quant.md) use FP4 values with two-level block scales, 4.25 bits per weight, and keep the LM head in int8 for accuracy. They cut the bytes per token by about a third and raise decode speed by 40% (Qwen3.5) to 45% (Qwen3, LFM2), at a measurable cost in perplexity that docs/quant.md reports per model.
The host is nearly out of the way. For LFM2 and Qwen3 the card runs one decode program compiled once, which reads the position from a register and looks up its own embedding and RoPE rows, and the logits stream back while the card is still running: the host adds 0.17 to 0.30 ms per token on omarchy (0.45 to 1.3 ms on opentpu). Qwen3.5 runs the same way for decode; its prefill still compiles each chunk's program on the host, ahead of the card.
Kernels in ol mlp, attention, full model layers
| @ol.jit
Language + compiler layouts, affine loop addressing, fusion
|
ISA 8 x 32-bit words per instruction
|
ISA simulator <======> RTL same bits, checked by the tests
(Python) (SystemVerilog)
| Vivado bitstream
FPGA card Kintex-7 xc7k480t
| PCIe
Host otpu-chat, otpu-smi, otpu-lens
The machine is deliberately simple. A sequencer issues one instruction per cycle to a few units: DMA moves data, the matrix unit multiplies int8 weights streamed from DRAM, the vector unit does fp32 math, and a quantizer turns results back into int8. There is no cache and no hidden scheduling: every data movement is an instruction, so a trace shows exactly where the cycles go. docs/isa.md describes the whole instruction set.
A kernel looks like this:
from opentpu import language as ol
@ol.jit
def mlp(h, gamma, w_gate, w_up, w_down, out, eps): # simplified; see kernels/mlp.py
x = ol.load(h)
xs = ol.quantize(rmsnorm(x, ol.load(gamma), eps))
g = ol.dot(xs, w_gate)
u = ol.dot(xs, w_up)
a = ol.all_gather(silu(g) * u)
y = ol.all_gather(ol.dot(a, w_down))
if ol.program_id() == 0:
ol.store(out, x + y)
Because every data movement is an instruction, a trace of a run explains its speed. Lens, the profiler, records a run from the RTL, the simulator or the card and opens it in the browser, with a roofline, a timeline and per-instruction tables (docs/lens.md).

Lens replaying part of a Qwen3 decode step. Colours show what each unit is doing in each cycle: busy, waiting on DRAM, or waiting on another instruction.
Everything except the card runs on a laptop.
pip install -e .
pip install pytest torch transformers
python3 -m pytest -q # RTL tests also need Verilator 5
hf download LiquidAI/LFM2.5-230M --local-dir models/LFM2.5-230M
otpu-chat --model lfm2 --backend isa # chat on the simulator
With a card, build the bitstream (make bit in boards/ypcb-00338), load
it over JTAG, then run sudo otpu-setup and otpu-chat --backend board.
docs/board.md walks through the bring-up.
| Command | What it does |
|---|---|
otpu-chat | chat with Qwen3-0.6B, LFM2.5-230M (--model lfm2), Qwen3.5-0.8B (--model qwen35), LFM2-2.6B (lfm2-2.6b), SmolLM3-3B (smollm3), Phi-4-mini (phi4-mini) or Qwen3.5-2B / 4B (qwen35-2b, qwen35-4b) |
otpu-smi | temperature, power, DRAM bandwidth and per-unit utilization |
otpu-lens | record a run and open it in the profiler |
otpu-selftest, otpu-diag | check that the card works |
opentpu/kernels and docs/compiler.md: how a kernel
becomes instructions.opentpu/isasim.py: the simulator, which is the spec.rtl/: the hardware, starting from rtl/top/otpu_top.sv.Issues and pull requests are welcome, and most of the work needs only Python and Verilator, not
an FPGA. Changes to the ISA, the simulator or the RTL must keep python3 -m pytest -q passing,
and performance claims should say how they were measured.
Apache License 2.0. See LICENSE.