RingForcing
Towards Precise Long-Term Memory for
Autoregressive Video Diffusion
When an object leaves the frame, does the same object reappear?
In autoregressive long-video generation models
When an object leaves the frame, does the same object reappear?



Ear


Memory capacity
How can we extend the effective context?
How can we extend the effective context?
Precise utilization
How can we precisely utilize long-range context?
Sliding Window
Limited context window.
Packed Context
Distant history is progressively coarsened.
Packed Context
Existing Methods

High noiseGlobal structure
Low noiseFine detail
Determines global semantic structure and motion dynamics.
Refines high-frequency textures, edges, and local details.
Preserve temporal continuity while coarsening spatial detail.
The context window now covers the full temporal span.
Preserve spatial detail through rotating sparse views.
Rotating offsets recover detail across the entire history.
The diffusion timestep selects how history is compressed.
Global history
Higher temporal densitySpatially coarse
Detail history
Spatially detailedTemporally sparse + rotating
Full historical coverage within a fixed token budget.
Memory capacity
✓ SolvedHow can we extend theeffective context?
Precise utilization
How can we preciselyutilize long-range context?




Adjacent frames are usually enough to predict what comes next.
Reverse + forward turns any video into a loop.
Shortcut removed.
AppearDisappearReappear
↻Clickany video to compare frames

