ngocdaobao/A-Survey-on-Looped-Transformers

114

5 commits

updated Sep 29, 2026

See the code

README

Looped Transformers: A Survey of Recurrent-Depth Architectures for Language Models

Contents


Taxonomy

Taxonomy of Looped Transformers in large language models

A. Loop Topology

A1. Whole Stack

The whole transformer stack/shared block is repeatedly applied.

A2. Sandwich: Prelude → Shared Core → Coda

Only the middle recurrent core is shared/looped.

A3. Layer-Local / Immediate

Each layer or small layer group is immediately repeated.

A4. Cyclic Patterns over M Unique Blocks

A set of unique blocks is reused following sequence/cycle patterns.

A5. Component-Level

Only a selected transformer component is recurrent.

A6. Hierarchical / Two-Timescale

Fast and slow recurrent modules operate at different timescales.

A7. Parallel / Pipelined across Loops

Loop computation is parallelized or pipelined.


B. Loop-Count Policy

B1. Fixed (T)

The model uses a fixed number of recurrent passes.

B2. Sampled at Train, Free at Test

Loop count varies during training and can be changed at inference.

B3. Supervised, Input-Dependent (T(n))

Loop count is supervised as a function of the input/problem size.

B4. Learned Online Halting

The stopping decision is made while recurrence is running.

B5. Learned Up-Front Routing

Depth is allocated before recurrent computation begins.

B6. Convergence Test / Training-Free Halting

The model stops when hidden states or outputs satisfy a convergence criterion.

B7. RL / Oracle-Supervised Stopping

Stopping policies are trained with reinforcement learning or oracle supervision.


C. Inter-Iteration State

C1. Pure Recurrence

Only the recurrent hidden state is passed forward.

C2. Input Injection / Recall

The original input is re-injected across recurrent steps.

C3. Predicted-Embedding Feedback

A predicted embedding is fed into the next recurrent iteration.

C4. Attention over Earlier Iterations

Later recurrent steps directly attend to earlier iterations.

C5. Fixed Point / Implicit

The recurrent computation is defined through an equilibrium/fixed point.


D. Sharing Strictness

D1. Exact Tie, No Conditioning

The same parameters are reused exactly across recurrent iterations.

D2. Timestep / Depth Encoding

The shared recurrent block is conditioned on the current loop/depth index.

D3. Per-Loop Scalars, Gates, or Norms

Small loop-specific parameters modulate the shared computation.

D4. Per-Loop Low-Rank Adapters

Iteration-specific low-rank adapters relax exact weight sharing.

D5. Experts Specializing per Pass

Sparse experts can specialize across recurrent iterations.


E. Training Recipe

E1. Provenance

From Scratch

Uptrained / Retrofitted

Training-Free

E2. Gradient Path

Full BPTT

Truncated BPTT

One-Step Approximation

Implicit Differentiation

Activation-Compressed BPTT

E3. Loop Schedule

Random Unrolling / Curriculum

Shortcut Consistency

  • LoopFormer — Ahmadreza Jeddi et al. arXiv, 2026.

Deep Supervision

RL over the Latent Trajectory

E4. Stability Machinery


F. Memory / KV Strategy

F1. Separate Cache per (Layer, Loop)

A separate KV cache is retained for every loop.

F2. Share First-Loop KV

KV states from the first recurrence are reused.

F3. Active-Token / Gated Single Cache

Only active token states are stored or one cache is updated with a gate.

F4. Compressed or Sub-Quadratic

The recurrent KV/attention representation is compressed or replaced by a more memory-efficient mechanism.

F5. Serving-Level Scheduling

The optimization is implemented at the inference/serving scheduler level.


Contributing

Contributions are welcome. Please keep new entries consistent with the format:

- **[Paper Title](paper-link)** — Author 1, Author 2, et al. *Venue, Year.*

ngocdaobao/A-Survey-on-Looped-Transformers

114

5 commits

updated Sep 29, 2026

See the code

README

Looped Transformers: A Survey of Recurrent-Depth Architectures for Language Models

Contents


Taxonomy

Taxonomy of Looped Transformers in large language models

A. Loop Topology

A1. Whole Stack

The whole transformer stack/shared block is repeatedly applied.

A2. Sandwich: Prelude → Shared Core → Coda

Only the middle recurrent core is shared/looped.

A3. Layer-Local / Immediate

Each layer or small layer group is immediately repeated.

A4. Cyclic Patterns over M Unique Blocks

A set of unique blocks is reused following sequence/cycle patterns.

A5. Component-Level

Only a selected transformer component is recurrent.

A6. Hierarchical / Two-Timescale

Fast and slow recurrent modules operate at different timescales.

A7. Parallel / Pipelined across Loops

Loop computation is parallelized or pipelined.


B. Loop-Count Policy

B1. Fixed (T)

The model uses a fixed number of recurrent passes.

B2. Sampled at Train, Free at Test

Loop count varies during training and can be changed at inference.

B3. Supervised, Input-Dependent (T(n))

Loop count is supervised as a function of the input/problem size.

B4. Learned Online Halting

The stopping decision is made while recurrence is running.

B5. Learned Up-Front Routing

Depth is allocated before recurrent computation begins.

B6. Convergence Test / Training-Free Halting

The model stops when hidden states or outputs satisfy a convergence criterion.

B7. RL / Oracle-Supervised Stopping

Stopping policies are trained with reinforcement learning or oracle supervision.


C. Inter-Iteration State

C1. Pure Recurrence

Only the recurrent hidden state is passed forward.

C2. Input Injection / Recall

The original input is re-injected across recurrent steps.

C3. Predicted-Embedding Feedback

A predicted embedding is fed into the next recurrent iteration.

C4. Attention over Earlier Iterations

Later recurrent steps directly attend to earlier iterations.

C5. Fixed Point / Implicit

The recurrent computation is defined through an equilibrium/fixed point.


D. Sharing Strictness

D1. Exact Tie, No Conditioning

The same parameters are reused exactly across recurrent iterations.

D2. Timestep / Depth Encoding

The shared recurrent block is conditioned on the current loop/depth index.

D3. Per-Loop Scalars, Gates, or Norms

Small loop-specific parameters modulate the shared computation.

D4. Per-Loop Low-Rank Adapters

Iteration-specific low-rank adapters relax exact weight sharing.

D5. Experts Specializing per Pass

Sparse experts can specialize across recurrent iterations.


E. Training Recipe

E1. Provenance

From Scratch

Uptrained / Retrofitted

Training-Free

E2. Gradient Path

Full BPTT

Truncated BPTT

One-Step Approximation

Implicit Differentiation

Activation-Compressed BPTT

E3. Loop Schedule

Random Unrolling / Curriculum

Shortcut Consistency

  • LoopFormer — Ahmadreza Jeddi et al. arXiv, 2026.

Deep Supervision

RL over the Latent Trajectory

E4. Stability Machinery


F. Memory / KV Strategy

F1. Separate Cache per (Layer, Loop)

A separate KV cache is retained for every loop.

F2. Share First-Loop KV

KV states from the first recurrence are reused.

F3. Active-Token / Gated Single Cache

Only active token states are stored or one cache is updated with a gate.

F4. Compressed or Sub-Quadratic

The recurrent KV/attention representation is compressed or replaced by a more memory-efficient mechanism.

F5. Serving-Level Scheduling

The optimization is implemented at the inference/serving scheduler level.


Contributing

Contributions are welcome. Please keep new entries consistent with the format:

- **[Paper Title](paper-link)** — Author 1, Author 2, et al. *Venue, Year.*