Residual-stream burden · Diffusion Transformers

What the residual stream must carry decides what later layers can learn

A Diffusion Transformer on large pixel patches trains well when it predicts the clean image and fails when it predicts the velocity. The two targets describe the same generative process. The difference is what the network has to keep in its residual stream until the readout.

clean image structure noise-dependent variation DiT block