The whole transformer stack/shared block is repeatedly applied.
Only the middle recurrent core is shared/looped.
Each layer or small layer group is immediately repeated.
A set of unique blocks is reused following sequence/cycle patterns.
Only a selected transformer component is recurrent.
Fast and slow recurrent modules operate at different timescales.
Loop computation is parallelized or pipelined.
The model uses a fixed number of recurrent passes.
Loop count varies during training and can be changed at inference.
Loop count is supervised as a function of the input/problem size.
The stopping decision is made while recurrence is running.
Depth is allocated before recurrent computation begins.
The model stops when hidden states or outputs satisfy a convergence criterion.
Stopping policies are trained with reinforcement learning or oracle supervision.
Only the recurrent hidden state is passed forward.
The original input is re-injected across recurrent steps.
A predicted embedding is fed into the next recurrent iteration.
Later recurrent steps directly attend to earlier iterations.
The recurrent computation is defined through an equilibrium/fixed point.
The same parameters are reused exactly across recurrent iterations.
The shared recurrent block is conditioned on the current loop/depth index.
Small loop-specific parameters modulate the shared computation.
Iteration-specific low-rank adapters relax exact weight sharing.
Sparse experts can specialize across recurrent iterations.
A separate KV cache is retained for every loop.
KV states from the first recurrence are reused.
Only active token states are stored or one cache is updated with a gate.
The recurrent KV/attention representation is compressed or replaced by a more memory-efficient mechanism.
The optimization is implemented at the inference/serving scheduler level.
Contributions are welcome. Please keep new entries consistent with the format:
- **[Paper Title](paper-link)** — Author 1, Author 2, et al. *Venue, Year.*
The whole transformer stack/shared block is repeatedly applied.
Only the middle recurrent core is shared/looped.
Each layer or small layer group is immediately repeated.
A set of unique blocks is reused following sequence/cycle patterns.
Only a selected transformer component is recurrent.
Fast and slow recurrent modules operate at different timescales.
Loop computation is parallelized or pipelined.
The model uses a fixed number of recurrent passes.
Loop count varies during training and can be changed at inference.
Loop count is supervised as a function of the input/problem size.
The stopping decision is made while recurrence is running.
Depth is allocated before recurrent computation begins.
The model stops when hidden states or outputs satisfy a convergence criterion.
Stopping policies are trained with reinforcement learning or oracle supervision.
Only the recurrent hidden state is passed forward.
The original input is re-injected across recurrent steps.
A predicted embedding is fed into the next recurrent iteration.
Later recurrent steps directly attend to earlier iterations.
The recurrent computation is defined through an equilibrium/fixed point.
The same parameters are reused exactly across recurrent iterations.
The shared recurrent block is conditioned on the current loop/depth index.
Small loop-specific parameters modulate the shared computation.
Iteration-specific low-rank adapters relax exact weight sharing.
Sparse experts can specialize across recurrent iterations.
A separate KV cache is retained for every loop.
KV states from the first recurrence are reused.
Only active token states are stored or one cache is updated with a gate.
The recurrent KV/attention representation is compressed or replaced by a more memory-efficient mechanism.
The optimization is implemented at the inference/serving scheduler level.
Contributions are welcome. Please keep new entries consistent with the format:
- **[Paper Title](paper-link)** — Author 1, Author 2, et al. *Venue, Year.*