sensenova/Looped-DiT-B32

Model

Looped-DiT B/32

3

1 commits

2 linked in READMEs

updated Sep 29, 2026

See the code

README

Looped-DiT B/32

Looped-DiT is a text-to-image diffusion transformer that runs a shared group of transformer blocks several times within each denoising step, so the model gets deeper without getting larger. Deep supervision trains the prediction after every loop, and exclusive self-attention (XSA) regulates the attention updates inside the loop. The backbone is the pixel-space MMDiT of MiniT2I, conditioned on a frozen FLAN-T5-Large.

This repository holds the EMA weights of Looped-DiT B/32 at step 290k:

  • patch size 32, 512 x 512 images
  • [6,5,6] blocks (pre-loop, looped, post-loop), with the 5 looped blocks run 4 times
  • trained with deep supervision, with XSA in the looped blocks

Results

100 Euler steps, classifier-free guidance 6.0, loop depth 4. TIIF-Short is TIIF-Bench scored on its short prompts.

GenEvalDPG-BenchPRISMCoReBenchSpatialGenEvalTIIF-ShortAvg
85.185.354.444.552.376.166.3

Usage

Download the checkpoint, then sample with the Looped-DiT code (link coming soon):

hf download sensenova/Looped-DiT-B32 looped-dit-b32.pt --local-dir checkpoints
python -m looped_dit.sample --checkpoint checkpoints/looped-dit-b32.pt \
    --prompt "a red cube on top of a blue sphere" --out sample.png

Add --loops 1 2 3 4 to render the prompt at several loop depths. Other depths work without retraining.

License

MIT

diffusion
looped-dit
pytorch
text-to-image

sensenova/Looped-DiT-B32

Model

Looped-DiT B/32

3

1 commits

2 linked in READMEs

updated Sep 29, 2026

See the code

README

Looped-DiT B/32

Looped-DiT is a text-to-image diffusion transformer that runs a shared group of transformer blocks several times within each denoising step, so the model gets deeper without getting larger. Deep supervision trains the prediction after every loop, and exclusive self-attention (XSA) regulates the attention updates inside the loop. The backbone is the pixel-space MMDiT of MiniT2I, conditioned on a frozen FLAN-T5-Large.

This repository holds the EMA weights of Looped-DiT B/32 at step 290k:

  • patch size 32, 512 x 512 images
  • [6,5,6] blocks (pre-loop, looped, post-loop), with the 5 looped blocks run 4 times
  • trained with deep supervision, with XSA in the looped blocks

Results

100 Euler steps, classifier-free guidance 6.0, loop depth 4. TIIF-Short is TIIF-Bench scored on its short prompts.

GenEvalDPG-BenchPRISMCoReBenchSpatialGenEvalTIIF-ShortAvg
85.185.354.444.552.376.166.3

Usage

Download the checkpoint, then sample with the Looped-DiT code (link coming soon):

hf download sensenova/Looped-DiT-B32 looped-dit-b32.pt --local-dir checkpoints
python -m looped_dit.sample --checkpoint checkpoints/looped-dit-b32.pt \
    --prompt "a red cube on top of a blue sphere" --out sample.png

Add --loops 1 2 3 4 to render the prompt at several loop depths. Other depths work without retraining.

License

MIT

diffusion
looped-dit
pytorch
text-to-image