What information is already available?
At denoising step i of video chunk j, ATC uses two cached hidden states. A hidden state is the diffusion transformer’s internal feature representation, before its output head converts it into a denoising velocity.
CURRENT CHUNK · EARLIER STEPa=hiaj The anchor a preserves the current chunk’s spatial layout, but comes from an earlier denoising stage (ia<i, usually ia=i−1).
PREVIOUS CHUNK · SAME STEPp=hij−1 The historical reference p is already at the target denoising stage, but motion may have shifted its content to different spatial positions.
The predictor also receives xemb, the current noisy latent after the input head, and C, the generation conditions and historical key–value (KV) context.
TRAINED FOR ITS OWN ROLLOUTS
Learn from the states
you will actually encounter.
Offline regression first teaches ATC to approximate the frozen model. Self-rollout on-policy distillation then adapts it to trajectories shaped by its own predictions, reducing the training–inference gap across both denoising steps and video chunks.
5.3%of a full DiT pass’s FLOPs
ATC on Self Forcing