Full-fabric VHDL Qwen3.5 9B/ Qwen3.8 27B LLM inference engine on XCVU33P/XCVU35P
VHDL
5
1,466 commits
updated Oct 4, 2026
Full-fabric VHDL LLM inference engine. Runs Qwen3.5-class transformer inference (9B on a single card, 27B targeted across two) entirely in FPGA fabric: INT4 streaming matvec, Gated DeltaNet, gated attention, and a transformer sequencer, with the weights resident in on-card HBM.
Outputs are validated against llama.cpp
running the BF16 GGUF. Its per-layer activations, captured through llama.cpp's
eval callback, are compared with a bit-accurate C model of the INT4 /
fixed-point datapath and with residuals read back from the card, and the
tokenizer and chat template are bit-exact against llama.cpp.
Details: tools/ref9b/README.md.
Target cards: SQRL FK33 (xcvu33p-fsvh2104-2L-e, 8 GiB HBM, PCIe Gen3 x4) and
Jungle Cat (2x xcvu35p modules, 8 GiB HBM each, Ethernet only, no PCIe).
Qwen3.5 is a hybrid model: most layers are a Gated DeltaNet linear-attention token mixer, the rest are gated full attention, and every layer ends in a SwiGLU MLP. The design splits into five subsystems (A-E); which mixer a given layer uses is decided per token by the sequencer's descriptor program, not wired in. Every matmul (projections, MLP, LM head) is served by the one INT4 matvec engine (subsystem A), and all state (INT4 weights, the KV cache, the GDN recurrent state, activations) lives in HBM.
flowchart TB
subgraph D["Subsystem D: transformer sequencer"]
seq["seq_desc_fetch + seq_opdec: per-token descriptor program"]
end
subgraph B["Subsystem B: Gated DeltaNet"]
gdn["gdn_block: causal conv, delta-rule recurrence, SiLU gate, QK L2-norm"]
end
subgraph C["Subsystem C: gated attention"]
attn["attn_block: RoPE, QK scores, softmax, score x V, output gate"]
end
host["Host: token ids in, argmax id out (PCIe Gen3 x4 / XDMA)"]
hbm[("HBM 8 GiB: INT4 weights, KV cache, GDN state, activations")]
n1["RMSNorm"]
mix{"layer type"}
r1["+ residual"]
n2["RMSNorm"]
mlp["SwiGLU MLP (swiglu_mem)"]
r2["+ residual"]
nf["final RMSNorm"]
lm["LM head"]
smp["sampler_stream: argmax"]
A["Subsystem A: INT4 streaming matvec (matvec_int4)"]
host --> seq
host <--> hbm
seq --> n1
n1 --> mix
mix -->|"linear-attn layers"| B
mix -->|"attention layers"| C
B --> r1
C --> r1
r1 --> n2
n2 --> mlp
mlp --> r2
r2 -->|"next layer"| n1
r2 -->|"after last layer"| nf
nf --> lm
lm --> smp
smp --> host
mlp -. matmuls .-> A
B -. matmuls .-> A
C -. matmuls .-> A
lm -. matmul .-> A
A <--> hbm
Subsystem E (the tensor-parallel collective for multi-die inference) is still a
skeleton; at NCARDS = 1 it is unused.
The shipping Qwen3.5 path, grouped by subsystem. rtl/ also contains
legacy stories260k / llama2-era units and *_skel.vhd skeletons that are
not part of this path, and several files are generated by tools/gen_*.py
(check line 2 before editing). Model shape lives in one place:
rtl/model_cfg_pkg.vhd.
| Subsystem | Top | Key units |
|---|---|---|
| A - INT4 streaming matvec (all matmuls) | rtl/matvec_int4.vhd | matvec_core (datapath), weight_streamer (INT4 weight/scale front end), act_mem_striped (banked activations), matvec_int4_desc_axi + matvec_int4_desc_pkg (descriptor-in-memory control), a_desc_adapter, a_job_counter |
| B - Gated DeltaNet (linear-attention layers) | rtl/gdn_block.vhd | gdn_conv (+gdn_conv_w_mem, gdn_conv_tap_mem) depthwise causal conv, gdn_recur_pipe delta-rule recurrence, gdn_scalar per-head scalar path, gdn_silu gate, l2norm_rs QK-norm, gdn_state_store (+gdn_state_mem/_axi, gdn_exp_mem/_capture) recurrent state in HBM, gdn_job_seq, gdn_emit_chain/_head_emit/_y_emit |
| C - gated attention (attention layers) | rtl/attn_block.vhd | attn_rope RoPE, attn_score_q12 QK scores, attn_softmax (+attn_recip), attn_mac_array score x V, attn_gate output gate, attn_kv_axi (+attn_kv_quant) KV cache in HBM, attn_emit, attn_twiddle, rmsnorm_rs input/QK norm |
| D - transformer sequencer | rtl/llama_top.vhd | seq_desc_fetch (per-token program from HBM), seq_opdec, seq_vec_issue/seq_vec_res, seq_region_lock + region_mem/region_drain (activation regions), rmsnorm_bf_mem top-level RMSNorm, swiglu_mem SwiGLU MLP, sampler_stream argmax |
| E - TP collective (future, multi-die) | rtl/tp_collective_skel.vhd | skeleton only |
| shared / HBM plumbing | - | axi_rd_port + axi_rd_fsm (HBM read masters), async_fifo, stream_fifo, bc_port_grant (B and C share one HBM read port) |
FK33 integration (hw/fk33/rtl/) | fk33_card.vhd | fk33_engine (subsystem A as a block-design cell), fk33_llama_top (B/C/D wrapper, generated), fk33_seam (host seam in front of D), fk33_eng_cdc (A clock-domain crossing) |
Several steps do direct DMA and MMIO to the card: reconfiguring the FPGA tears
the PCIe bus down and rescans it, and the weight load writes into HBM. An
interrupted or faulted transaction can hang the host, so run these steps
interactively and watch them, not from cron, CI, or an unattended script. See
docs/2026-08-27_fk33-pcie-bringup-procedure.md for the long form.
Prerequisites: one or two FK33s in PCIe slots, Vivado 2023.2, a Linux host,
and a Qwen3.5-9B INT4 weight image with its MANIFEST.json (layout in
docs/2026-08-27_hbm-residency-map.md).
Common to both modes, build the bitstream and set up the host driver:
# Build the endpoint bitstream (~1-1.5 h; or reuse hw/fk33/bit/*.bit).
# FK33_CARD=1 adds the B/C/D subsystems; engine-only omits it.
FK33_CARD=1 hw/fk33/pcieep_build.sh
# XDMA driver (once), then group access to /dev/xdma* (once per boot).
hw/fk33/host/build_xdma_driver.sh # compiles, installs nothing
sudo hw/fk33/host/setup-fk33-access.sh
make -C server libqwen35chat.so # tokenizer + chat template shim
There are two ways to deploy.
One FK33 holds the whole 9B: every block, the LM head, the KV cache and the GDN state in its 8 GiB of HBM.
# Program the card over JTAG (FK33_CARD=<JTAG serial> picks it).
hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
# Load the full 9B image and verify it (defaults to /dev/xdma0).
hw/fk33/host/fk33_load_weights.py load MANIFEST.json --verify
# Sanity check, gate on CTXTEST_PASS.
hw/fk33/host/fk33_ctxtest.sh single <outdir> 256
# Serve the web UI.
python3 server/llmvhdl_server.py --mode single
The model is split as a pipeline: card 0 holds blocks 0..k-1 and ends with
the residual, card 1 holds blocks k..N-1 plus the LM head, and the host
carries the residual between them each token. This roughly doubles decode
throughput over one card (about 2.5 tokens/s measured). Each card gets its own
image with its own block range; the split point is read off the manifests, not
configured anywhere else.
# Program both cards with the same bitstream, one at a time.
FK33_CARD=<serial A> hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
FK33_CARD=<serial B> hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
# Load each card's half of the image into its own node.
FK33_H2C=/dev/xdma0_h2c_0 FK33_C2H=/dev/xdma0_c2h_0 \
hw/fk33/host/fk33_load_weights.py load MANIFEST.first.json --verify
FK33_H2C=/dev/xdma1_h2c_0 FK33_C2H=/dev/xdma1_c2h_0 \
hw/fk33/host/fk33_load_weights.py load MANIFEST.second.json --verify
# Sanity check the pair, gate on CTXTEST_PASS.
hw/fk33/host/fk33_ctxtest.sh pair <outdir> 256
# Serve the web UI (pair is the default mode).
python3 server/llmvhdl_server.py
The xdma0/xdma1 node numbers can swap on every reload, so the host assigns
each card's role from the image it holds (the card with the LM head is card 1),
not from the node number.
The cards have been tested on silicon at 75 MHz only; nothing above 75 MHz, and
nothing at the full 0.85 V VCCINT, has been run on hardware yet. Getting to
higher clocks is ongoing work. The standalone per-block ratings below (each block
placed and routed alone with tools/rate/, full table in hw/targets/ratings/)
suggest there is headroom on both parts:
| block | what it is | FK33 (vu33p -2LV, 0.72 V) | VU35P (-2, 0.85 V) |
|---|---|---|---|
c_attn | attention MAC array | 150.4 | 204.5 |
c_kv | KV cache over HBM | 153.8 | 228.2 |
v_swg | SwiGLU | 183.5 | 242.9 |
a_engine | INT4 matvec | 185.5 | 264.0 |
b_gdn | Gated DeltaNet | 194.9 | 254.3 |
v_rms | RMSNorm | 206.9 | 281.9 |
Either way, open http://127.0.0.1:8000/ for the chat UI (a copy of the
llama.cpp web UI). It is greedy-only, one request at a time, with no prompt
cache; the docstring in server/llmvhdl_server.py explains why.
The two-die 27B target runs on the SQRL Jungle Cat (JCC2L-Lite carrier, two
JCM35P modules = 2x xcvu35p, -2L). 27B INT4 fits one VU35P at about 46% LUT,
and every shipping block closes 200 MHz on the VU35P -2 grade standalone.
Working: both modules program and come up with a standard .bit over
Ethernet through SQRL's on-board STM32, which bridges to JTAG (sqrl_bridge
CoE; Vivado drives it over XVC). IDCODE confirmed on both dies. No PCIe and no
bitstream reverse-engineering required.
Blockers:
See docs/boards/jungle-cat/2026-09-27_bringup.md and
docs/2026-09-24_jungle-cat-performance-estimate.md.
docs/WORKLOG.md is the live board; docs/ holds the dated design and
debugging notes. In flight:
docs/PLAN_TO_FIRST_INFERENCE.md and the two-card pipeline spec.docs/2026-09-24_jungle-cat-performance-estimate.md.hw/targets/.MIT (see LICENSE). Third-party components and their licenses are
listed in NOTICE; the web UI in ui/ is derived from the llama.cpp
web UI (MIT) and keeps its upstream notice in ui/LICENSE.llama.cpp.
VHDL
49.2%
Python
15.2%
Shell
10.7%
C
8.9%
TypeScript
5.9%
Tcl
5.6%
Svelte
3.7%
Full-fabric VHDL Qwen3.5 9B/ Qwen3.8 27B LLM inference engine on XCVU33P/XCVU35P
VHDL
5
1,466 commits
updated Oct 4, 2026
Full-fabric VHDL LLM inference engine. Runs Qwen3.5-class transformer inference (9B on a single card, 27B targeted across two) entirely in FPGA fabric: INT4 streaming matvec, Gated DeltaNet, gated attention, and a transformer sequencer, with the weights resident in on-card HBM.
Outputs are validated against llama.cpp
running the BF16 GGUF. Its per-layer activations, captured through llama.cpp's
eval callback, are compared with a bit-accurate C model of the INT4 /
fixed-point datapath and with residuals read back from the card, and the
tokenizer and chat template are bit-exact against llama.cpp.
Details: tools/ref9b/README.md.
Target cards: SQRL FK33 (xcvu33p-fsvh2104-2L-e, 8 GiB HBM, PCIe Gen3 x4) and
Jungle Cat (2x xcvu35p modules, 8 GiB HBM each, Ethernet only, no PCIe).
Qwen3.5 is a hybrid model: most layers are a Gated DeltaNet linear-attention token mixer, the rest are gated full attention, and every layer ends in a SwiGLU MLP. The design splits into five subsystems (A-E); which mixer a given layer uses is decided per token by the sequencer's descriptor program, not wired in. Every matmul (projections, MLP, LM head) is served by the one INT4 matvec engine (subsystem A), and all state (INT4 weights, the KV cache, the GDN recurrent state, activations) lives in HBM.
flowchart TB
subgraph D["Subsystem D: transformer sequencer"]
seq["seq_desc_fetch + seq_opdec: per-token descriptor program"]
end
subgraph B["Subsystem B: Gated DeltaNet"]
gdn["gdn_block: causal conv, delta-rule recurrence, SiLU gate, QK L2-norm"]
end
subgraph C["Subsystem C: gated attention"]
attn["attn_block: RoPE, QK scores, softmax, score x V, output gate"]
end
host["Host: token ids in, argmax id out (PCIe Gen3 x4 / XDMA)"]
hbm[("HBM 8 GiB: INT4 weights, KV cache, GDN state, activations")]
n1["RMSNorm"]
mix{"layer type"}
r1["+ residual"]
n2["RMSNorm"]
mlp["SwiGLU MLP (swiglu_mem)"]
r2["+ residual"]
nf["final RMSNorm"]
lm["LM head"]
smp["sampler_stream: argmax"]
A["Subsystem A: INT4 streaming matvec (matvec_int4)"]
host --> seq
host <--> hbm
seq --> n1
n1 --> mix
mix -->|"linear-attn layers"| B
mix -->|"attention layers"| C
B --> r1
C --> r1
r1 --> n2
n2 --> mlp
mlp --> r2
r2 -->|"next layer"| n1
r2 -->|"after last layer"| nf
nf --> lm
lm --> smp
smp --> host
mlp -. matmuls .-> A
B -. matmuls .-> A
C -. matmuls .-> A
lm -. matmul .-> A
A <--> hbm
Subsystem E (the tensor-parallel collective for multi-die inference) is still a
skeleton; at NCARDS = 1 it is unused.
The shipping Qwen3.5 path, grouped by subsystem. rtl/ also contains
legacy stories260k / llama2-era units and *_skel.vhd skeletons that are
not part of this path, and several files are generated by tools/gen_*.py
(check line 2 before editing). Model shape lives in one place:
rtl/model_cfg_pkg.vhd.
| Subsystem | Top | Key units |
|---|---|---|
| A - INT4 streaming matvec (all matmuls) | rtl/matvec_int4.vhd | matvec_core (datapath), weight_streamer (INT4 weight/scale front end), act_mem_striped (banked activations), matvec_int4_desc_axi + matvec_int4_desc_pkg (descriptor-in-memory control), a_desc_adapter, a_job_counter |
| B - Gated DeltaNet (linear-attention layers) | rtl/gdn_block.vhd | gdn_conv (+gdn_conv_w_mem, gdn_conv_tap_mem) depthwise causal conv, gdn_recur_pipe delta-rule recurrence, gdn_scalar per-head scalar path, gdn_silu gate, l2norm_rs QK-norm, gdn_state_store (+gdn_state_mem/_axi, gdn_exp_mem/_capture) recurrent state in HBM, gdn_job_seq, gdn_emit_chain/_head_emit/_y_emit |
| C - gated attention (attention layers) | rtl/attn_block.vhd | attn_rope RoPE, attn_score_q12 QK scores, attn_softmax (+attn_recip), attn_mac_array score x V, attn_gate output gate, attn_kv_axi (+attn_kv_quant) KV cache in HBM, attn_emit, attn_twiddle, rmsnorm_rs input/QK norm |
| D - transformer sequencer | rtl/llama_top.vhd | seq_desc_fetch (per-token program from HBM), seq_opdec, seq_vec_issue/seq_vec_res, seq_region_lock + region_mem/region_drain (activation regions), rmsnorm_bf_mem top-level RMSNorm, swiglu_mem SwiGLU MLP, sampler_stream argmax |
| E - TP collective (future, multi-die) | rtl/tp_collective_skel.vhd | skeleton only |
| shared / HBM plumbing | - | axi_rd_port + axi_rd_fsm (HBM read masters), async_fifo, stream_fifo, bc_port_grant (B and C share one HBM read port) |
FK33 integration (hw/fk33/rtl/) | fk33_card.vhd | fk33_engine (subsystem A as a block-design cell), fk33_llama_top (B/C/D wrapper, generated), fk33_seam (host seam in front of D), fk33_eng_cdc (A clock-domain crossing) |
Several steps do direct DMA and MMIO to the card: reconfiguring the FPGA tears
the PCIe bus down and rescans it, and the weight load writes into HBM. An
interrupted or faulted transaction can hang the host, so run these steps
interactively and watch them, not from cron, CI, or an unattended script. See
docs/2026-08-27_fk33-pcie-bringup-procedure.md for the long form.
Prerequisites: one or two FK33s in PCIe slots, Vivado 2023.2, a Linux host,
and a Qwen3.5-9B INT4 weight image with its MANIFEST.json (layout in
docs/2026-08-27_hbm-residency-map.md).
Common to both modes, build the bitstream and set up the host driver:
# Build the endpoint bitstream (~1-1.5 h; or reuse hw/fk33/bit/*.bit).
# FK33_CARD=1 adds the B/C/D subsystems; engine-only omits it.
FK33_CARD=1 hw/fk33/pcieep_build.sh
# XDMA driver (once), then group access to /dev/xdma* (once per boot).
hw/fk33/host/build_xdma_driver.sh # compiles, installs nothing
sudo hw/fk33/host/setup-fk33-access.sh
make -C server libqwen35chat.so # tokenizer + chat template shim
There are two ways to deploy.
One FK33 holds the whole 9B: every block, the LM head, the KV cache and the GDN state in its 8 GiB of HBM.
# Program the card over JTAG (FK33_CARD=<JTAG serial> picks it).
hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
# Load the full 9B image and verify it (defaults to /dev/xdma0).
hw/fk33/host/fk33_load_weights.py load MANIFEST.json --verify
# Sanity check, gate on CTXTEST_PASS.
hw/fk33/host/fk33_ctxtest.sh single <outdir> 256
# Serve the web UI.
python3 server/llmvhdl_server.py --mode single
The model is split as a pipeline: card 0 holds blocks 0..k-1 and ends with
the residual, card 1 holds blocks k..N-1 plus the LM head, and the host
carries the residual between them each token. This roughly doubles decode
throughput over one card (about 2.5 tokens/s measured). Each card gets its own
image with its own block range; the split point is read off the manifests, not
configured anywhere else.
# Program both cards with the same bitstream, one at a time.
FK33_CARD=<serial A> hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
FK33_CARD=<serial B> hw/fk33/host/fk33_reload.sh hw/fk33/bit/fk33_pcieep_eng.bit
# Load each card's half of the image into its own node.
FK33_H2C=/dev/xdma0_h2c_0 FK33_C2H=/dev/xdma0_c2h_0 \
hw/fk33/host/fk33_load_weights.py load MANIFEST.first.json --verify
FK33_H2C=/dev/xdma1_h2c_0 FK33_C2H=/dev/xdma1_c2h_0 \
hw/fk33/host/fk33_load_weights.py load MANIFEST.second.json --verify
# Sanity check the pair, gate on CTXTEST_PASS.
hw/fk33/host/fk33_ctxtest.sh pair <outdir> 256
# Serve the web UI (pair is the default mode).
python3 server/llmvhdl_server.py
The xdma0/xdma1 node numbers can swap on every reload, so the host assigns
each card's role from the image it holds (the card with the LM head is card 1),
not from the node number.
The cards have been tested on silicon at 75 MHz only; nothing above 75 MHz, and
nothing at the full 0.85 V VCCINT, has been run on hardware yet. Getting to
higher clocks is ongoing work. The standalone per-block ratings below (each block
placed and routed alone with tools/rate/, full table in hw/targets/ratings/)
suggest there is headroom on both parts:
| block | what it is | FK33 (vu33p -2LV, 0.72 V) | VU35P (-2, 0.85 V) |
|---|---|---|---|
c_attn | attention MAC array | 150.4 | 204.5 |
c_kv | KV cache over HBM | 153.8 | 228.2 |
v_swg | SwiGLU | 183.5 | 242.9 |
a_engine | INT4 matvec | 185.5 | 264.0 |
b_gdn | Gated DeltaNet | 194.9 | 254.3 |
v_rms | RMSNorm | 206.9 | 281.9 |
Either way, open http://127.0.0.1:8000/ for the chat UI (a copy of the
llama.cpp web UI). It is greedy-only, one request at a time, with no prompt
cache; the docstring in server/llmvhdl_server.py explains why.
The two-die 27B target runs on the SQRL Jungle Cat (JCC2L-Lite carrier, two
JCM35P modules = 2x xcvu35p, -2L). 27B INT4 fits one VU35P at about 46% LUT,
and every shipping block closes 200 MHz on the VU35P -2 grade standalone.
Working: both modules program and come up with a standard .bit over
Ethernet through SQRL's on-board STM32, which bridges to JTAG (sqrl_bridge
CoE; Vivado drives it over XVC). IDCODE confirmed on both dies. No PCIe and no
bitstream reverse-engineering required.
Blockers:
See docs/boards/jungle-cat/2026-09-27_bringup.md and
docs/2026-09-24_jungle-cat-performance-estimate.md.
docs/WORKLOG.md is the live board; docs/ holds the dated design and
debugging notes. In flight:
docs/PLAN_TO_FIRST_INFERENCE.md and the two-card pipeline spec.docs/2026-09-24_jungle-cat-performance-estimate.md.hw/targets/.MIT (see LICENSE). Third-party components and their licenses are
listed in NOTICE; the web UI in ui/ is derived from the llama.cpp
web UI (MIT) and keeps its upstream notice in ui/LICENSE.llama.cpp.
VHDL
49.2%
Python
15.2%
Shell
10.7%
C
8.9%
TypeScript
5.9%
Tcl
5.6%
Svelte
3.7%