ECE Seminar Lecture Series

Demystifying Manifold Constraints in LLM Pre-Training

Shiqian Ma is a Professor in the Data Science and AI Institute, the Department of Applied Mathematics and Statistics, the Department of Computer Science and the Department of Electrical and Computer Engineering at Johns Hopkins University.

Wednesday, September 30, 2026
Noon–1 p.m.

601 Computer Studies Building

Man smiling at camera wearing glasses.Abstract: The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While recent constrained optimization approaches that explicitly restrict weights may improve numerical stability and performance, the mechanism and motivation for adding constraints still remain elusive. In this work, we systematically demystify the role of explicit manifold constraints in LLM pre-training. By introducing the Msign-Aligned Constrained Riemannian Optimizer (MACRO) -- a provably convergent, single-loop optimization framework -- our study disentangles weight regularization heuristics from interacting mechanisms like RMS normalization and decoupled weight decay. Theoretical analyses and comprehensive empirical evaluations reveal that manifold constraints independently bound forward activation scales and enforce stable rotational equilibrium, thereby subsuming the roles of these heuristic mechanisms. Evaluations on large-scale LLM architectures demonstrate that MACRO achieves highly competitive performance.  

Bio: Shiqian Ma is a Professor in the Data Science and AI Institute, the Department of Applied Mathematics and Statistics, the Department of Computer Science and the Department of Electrical and Computer Engineering at Johns Hopkins University. His research focuses on the mathematical foundations of modern AI, with particular emphasis on optimization, geometry, and machine learning. He develops scalable optimization algorithms and mathematical frameworks for efficient and reliable training of foundation models and large-scale machine learning systems. Ma currently serves as an Action Editor for the Journal of Machine Learning Research (JMLR) and Transactions on Machine Learning Research (TMLR), and as an Associate Editor for SIAM Journal on Optimization (SIOPT) and Journal of Optimization Theory and Applications (JOTA). He has served as a Senior Area Chair for NeurIPS, ICML, and AISTATS, among other leading machine learning conferences. His research is supported by the National Science Foundation and the Office of Naval Research. His work on Riemannian optimization has been recognized by the 2024 INFORMS Computing Society Prize and the 2024 SIAM Review SIGEST Award.