What the residual stream must carry decides what later layers can learn
A Diffusion Transformer on large pixel patches trains well when it predicts the clean image and fails when it predicts the velocity. The two targets describe the same generative process. The difference is what the network has to keep in its residual stream until the readout.