Veda Sparse Attention for ComfyUI
See the codePaper · Project Page · Predictor · Training Code · 中文说明
Veda is a learned sparse-attention method for video diffusion models. Attention dominates the cost of a long clip, and a small distilled predictor can say in advance which tiles of the attention map carry the result. Veda computes roughly the top 10% of them and skips the rest.
This repository packages Veda as a custom node for ComfyUI's native MiniMax-H3 (T2VA, FL2VA and R2VA): one node, MODEL in and MODEL out.
Requires ComfyUI >= 0.38.0. In ComfyUI Manager, search "Veda". Or:
comfy node install veda-sparse-attention
Or clone this repository into ComfyUI/custom_nodes/ and restart ComfyUI.
Open Workflow -> Browse Templates -> Veda-on-ComfyUI and pick "Veda MiniMax H3 T2VA" or "Veda MiniMax H3 R2VA". The missing-model dialog offers the predictor. Write a prompt and run.
To add Veda to an existing H3 workflow, put Veda Sparse Attention (MiniMax H3) on the MODEL wire after the model and any LoRA loaders, last before the guider or sampler. Select it and press Ctrl+B to bypass it and render the same seed with full attention.
Do not combine it with ComfyUI's own "Model Sparse Attention" node on H3: that one replaces the attention blocks outright, so Veda would never be called. The node says so when it sees both.
Other attention nodes can stay. Veda runs on top of "Patch Sage Attention KJ" (the calls it does not take go to sage), and KJNodes' "MiniMax H3 Low VRAM Attention" keeps working. A node that swaps out the whole H3 attention, such as "MiniMax H3 Mem Eff Sage Attention Patch", must come before Veda: Veda then runs the sparse layers and leaves that node the full-attention ones, and says so on its node. Placed after Veda it would hide every call from Veda; the node warns about that too.
The node does not download anything itself. The predictor (275 MB)
reaches models/veda one of two ways:
minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors from the
predictor repository
into ComfyUI/models/veda/ and restart ComfyUI.Any .safetensors in that folder appears in the predictor list, so a
predictor trained elsewhere is selected the same way. If the selected
file is not there, the node says so and prints the URL rather than
reaching out on its own.
Veda done · Triton INT8 (SM120)
Video: 1344x768 · 5.2 s
Attention computed: 10.9% of full attention (89.1% skipped)
Before sampling the node names the kernel it will use and the sparsity;
while sampling, the video size and the trained tile plan it matched; at
the end, how much of full attention was actually computed. A line
starting Veda off gives the reason Veda is not running, and nearest trained size: ... on the tile plan line means the video is outside the
trained set. verbose adds per-phase timing, call counts and predictor
details.
Warnings and errors come in English and Chinese, one after the other; the node's name and tooltips follow ComfyUI's interface language.
Only model and predictor are visible. The rest are advanced inputs
(click "show advanced inputs") and default to the trained values.
| Input | Default | Description |
|---|---|---|
generated_sparsity | 90% | Sparsity of the generated video's attention. 90% skips 90% of the key tiles each query tile could attend, which is the trained value; lower is closer to full attention and slower. A whole number such as 24 keeps exactly that many 128-token key tiles instead. |
reference_sparsity | 90% | The same for references: first/last frames, guide frames, reference images and videos. 0% gives them full attention. |
full_attention_layers | empty | 0-based DiT blocks that keep full attention, e.g. 0, 1, 47-49. |
full_attention_steps | empty | 0-based sampling steps that keep full attention, e.g. 0. |
verbose | off | Also report attention time per phase, call counts and predictor details on the node after each run. |
The released predictor was trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5 / 10 / 14 s with the 8-step Turbo LoRA. Other sizes fall back to the tile plan of the nearest aspect ratio and duration. Other sizes, other step counts and R2VA / FL2VA references all work, but are outside the training data, so compare them against full attention.
T2VA at 1344x768, 124 frames (5.2 s), 8-step Turbo LoRA, 90% sparsity, on an RTX 5070 12 GB under Windows 11.
| Attention | Per step | 8 steps | Attention per step |
|---|---|---|---|
| ComfyUI default | 40.7 s | 342 s | 31.1 s |
ComfyUI --use-sage-attention | 24.8 s | 231 s | 15.2 s |
| Veda sparse INT8 | 14.0 s | 130 s | 4.41 s |
One attention layer at 104k tokens and 90% sparsity takes 528 ms against
11.8 s for full attention. Veda's INT8 arithmetic is the same code
ComfyUI runs behind --use-sage-attention, so the quality is what that
path gives, with sparsity on top. What remains per step is MLP and weight
movement, not attention.
One Triton kernel covers every NVIDIA GPU from SM80 on, Windows and Linux alike. Older cards, ROCm and CPU have no kernel: the node says so and the model runs its own attention.
| Hardware | Status |
|---|---|
| RTX 30 / A100 / RTX 40 / L40 (sm80-89) | code path ready, not yet verified on this hardware |
| H100 / H200 (sm90) | code path ready, not yet verified on this hardware |
| B200 / B300 (sm100 / sm103) | code path ready, not yet verified on this hardware |
| RTX 50, RTX PRO 6000 Blackwell (sm120) | verified: RTX 5070, Windows 11 |
| DGX Spark / GB10 (sm121) | code path ready, not yet verified on this hardware |
| Apple silicon (M series) | verified: M3 Pro, macOS 15 |
Per-architecture notes and the full measurement log are in docs/hardware.md.
Code: MIT. The INT8 kernel's arithmetic is derived from SageAttention v1 (BSD-3-Clause). The predictor inherits the MiniMax H3 Community License from its base model. See NOTICE.md.
@inproceedings{han2026veda,
title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
and Jiang, Yi and Qi, Xiaojuan},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
Veda Sparse Attention for ComfyUI
See the codePaper · Project Page · Predictor · Training Code · 中文说明
Veda is a learned sparse-attention method for video diffusion models. Attention dominates the cost of a long clip, and a small distilled predictor can say in advance which tiles of the attention map carry the result. Veda computes roughly the top 10% of them and skips the rest.
This repository packages Veda as a custom node for ComfyUI's native MiniMax-H3 (T2VA, FL2VA and R2VA): one node, MODEL in and MODEL out.
Requires ComfyUI >= 0.38.0. In ComfyUI Manager, search "Veda". Or:
comfy node install veda-sparse-attention
Or clone this repository into ComfyUI/custom_nodes/ and restart ComfyUI.
Open Workflow -> Browse Templates -> Veda-on-ComfyUI and pick "Veda MiniMax H3 T2VA" or "Veda MiniMax H3 R2VA". The missing-model dialog offers the predictor. Write a prompt and run.
To add Veda to an existing H3 workflow, put Veda Sparse Attention (MiniMax H3) on the MODEL wire after the model and any LoRA loaders, last before the guider or sampler. Select it and press Ctrl+B to bypass it and render the same seed with full attention.
Do not combine it with ComfyUI's own "Model Sparse Attention" node on H3: that one replaces the attention blocks outright, so Veda would never be called. The node says so when it sees both.
Other attention nodes can stay. Veda runs on top of "Patch Sage Attention KJ" (the calls it does not take go to sage), and KJNodes' "MiniMax H3 Low VRAM Attention" keeps working. A node that swaps out the whole H3 attention, such as "MiniMax H3 Mem Eff Sage Attention Patch", must come before Veda: Veda then runs the sparse layers and leaves that node the full-attention ones, and says so on its node. Placed after Veda it would hide every call from Veda; the node warns about that too.
The node does not download anything itself. The predictor (275 MB)
reaches models/veda one of two ways:
minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors from the
predictor repository
into ComfyUI/models/veda/ and restart ComfyUI.Any .safetensors in that folder appears in the predictor list, so a
predictor trained elsewhere is selected the same way. If the selected
file is not there, the node says so and prints the URL rather than
reaching out on its own.
Veda done · Triton INT8 (SM120)
Video: 1344x768 · 5.2 s
Attention computed: 10.9% of full attention (89.1% skipped)
Before sampling the node names the kernel it will use and the sparsity;
while sampling, the video size and the trained tile plan it matched; at
the end, how much of full attention was actually computed. A line
starting Veda off gives the reason Veda is not running, and nearest trained size: ... on the tile plan line means the video is outside the
trained set. verbose adds per-phase timing, call counts and predictor
details.
Warnings and errors come in English and Chinese, one after the other; the node's name and tooltips follow ComfyUI's interface language.
Only model and predictor are visible. The rest are advanced inputs
(click "show advanced inputs") and default to the trained values.
| Input | Default | Description |
|---|---|---|
generated_sparsity | 90% | Sparsity of the generated video's attention. 90% skips 90% of the key tiles each query tile could attend, which is the trained value; lower is closer to full attention and slower. A whole number such as 24 keeps exactly that many 128-token key tiles instead. |
reference_sparsity | 90% | The same for references: first/last frames, guide frames, reference images and videos. 0% gives them full attention. |
full_attention_layers | empty | 0-based DiT blocks that keep full attention, e.g. 0, 1, 47-49. |
full_attention_steps | empty | 0-based sampling steps that keep full attention, e.g. 0. |
verbose | off | Also report attention time per phase, call counts and predictor details on the node after each run. |
The released predictor was trained for 1344x768, 768x1344, 768x768 and 1024x768 at 5 / 10 / 14 s with the 8-step Turbo LoRA. Other sizes fall back to the tile plan of the nearest aspect ratio and duration. Other sizes, other step counts and R2VA / FL2VA references all work, but are outside the training data, so compare them against full attention.
T2VA at 1344x768, 124 frames (5.2 s), 8-step Turbo LoRA, 90% sparsity, on an RTX 5070 12 GB under Windows 11.
| Attention | Per step | 8 steps | Attention per step |
|---|---|---|---|
| ComfyUI default | 40.7 s | 342 s | 31.1 s |
ComfyUI --use-sage-attention | 24.8 s | 231 s | 15.2 s |
| Veda sparse INT8 | 14.0 s | 130 s | 4.41 s |
One attention layer at 104k tokens and 90% sparsity takes 528 ms against
11.8 s for full attention. Veda's INT8 arithmetic is the same code
ComfyUI runs behind --use-sage-attention, so the quality is what that
path gives, with sparsity on top. What remains per step is MLP and weight
movement, not attention.
One Triton kernel covers every NVIDIA GPU from SM80 on, Windows and Linux alike. Older cards, ROCm and CPU have no kernel: the node says so and the model runs its own attention.
| Hardware | Status |
|---|---|
| RTX 30 / A100 / RTX 40 / L40 (sm80-89) | code path ready, not yet verified on this hardware |
| H100 / H200 (sm90) | code path ready, not yet verified on this hardware |
| B200 / B300 (sm100 / sm103) | code path ready, not yet verified on this hardware |
| RTX 50, RTX PRO 6000 Blackwell (sm120) | verified: RTX 5070, Windows 11 |
| DGX Spark / GB10 (sm121) | code path ready, not yet verified on this hardware |
| Apple silicon (M series) | verified: M3 Pro, macOS 15 |
Per-architecture notes and the full measurement log are in docs/hardware.md.
Code: MIT. The INT8 kernel's arithmetic is derived from SageAttention v1 (BSD-3-Clause). The predictor inherits the MiniMax H3 Community License from its base model. See NOTICE.md.
@inproceedings{han2026veda,
title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
and Jiang, Yi and Qi, Xiaojuan},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}