About Me

I am a PhD student in the Department of Mathematics at UC San Diego (since 2023), working with Prof. Alex Cloninger, Prof. Rahul Parhi, and Prof. Yu-Xiang Wang on deep learning research. I received both my B.S. and M.S. degrees in Mathematics from Southern University of Science and Technology, where I was advised by Prof. Yifei Zhu.

My research focuses on understanding how training organizes information from data into parameterized computational structures. My goal is to use this understanding to develop better neural architectures and training methods for foundation models.

My work on neural shattering, data geometry, and sparse connectivity examines how data geometry and network structure shape the implicit bias of gradient descent. Extending this perspective to the design of generative models, my recent work on residual-stream burden investigates how prediction objectives and architecture jointly shape the representations learned by diffusion Transformers, and turns these findings into concrete architectural improvements.

Papers

Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
Tongtong Liang, Siqi Kou, Ziqiao Xi, Esha Singh, Kun Zhou, Zhijie Deng, Alexander Cloninger, Yu-Xiang Wang, Rahul Parhi
Preprint · 2026 · arXiv · Project page · Code

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Zhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong, Songchun Zhang, Junchao Huang, Guian Fang, Xin Wang, Pan Hui, Anyi Rao
Preprint · 2026 · arXiv · Project page · Code

What Architectural Inductive Bias Makes Diffusion Models Succeed? A Perspective from the Implicit Regularization of Gradient Descent
Tongtong Liang, Esha Singh, Rahul Parhi, Alexander Cloninger, Yu-Xiang Wang
ICML 2026 Workshop on Foundations of Deep Generative Models · FoGen 2026

Does Sparse Connectivity Improve Generalization? Convolutional Networks Below the Edge of Stability
Tongtong Liang, Esha Singh, Rahul Parhi, Alexander Cloninger, Yu-Xiang Wang
NeurIPS 2026 · arXiv

IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
Zhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur, Shibo Hao, Varad Pimpalkhute, Tongtong Liang, Feng Yao, Zhengzhong Liu, Eric Xing, Virginia Smith, Ruslan Salakhutdinov, Zhiting Hu, Taylor Killian, Aviral Kumar
ICML 2026 · arXiv · blog

Generalization Below the Edge of Stability: The Role of Data Geometry
Tongtong Liang, Alexander Cloninger, Rahul Parhi, Yu-Xiang Wang
ICLR 2026 · arXiv

Stable Minima of ReLU Neural Networks Suffer from the Curse of Dimensionality: The Neural Shattering Phenomenon
Tongtong Liang, Dan Qiao, Yu-Xiang Wang, Rahul Parhi
NeurIPS 2025 Spotlight · arXiv