Publications

2026

  1. arXiv
    Muon^2: Boosting Muon via Adaptive Second-Moment Preconditioning
    Ziyue Liu, Ruijie Zhang, Zhengyang Wang, and 4 more authors
    arXiv preprint arXiv:2604.09967, 2026
  2. arXiv
    Muon+: Towards better muon via one additional normalization step
    Ruijie Zhang, Yequan Zhao, Ziyue Liu, and 2 more authors
    arXiv e-prints, 2026
  3. arXiv
    TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training
    Ruijie Zhang, Yequan Zhao, Ziyue Liu, and 5 more authors
    arXiv preprint arXiv:2601.23261, 2026
  4. arXiv
    ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
    Ziyue Liu, Zhengyang Wang, Ruijie Zhang, and 7 more authors
    arXiv preprint arXiv:2605.11215, 2026

2025

  1. EMNLP
    Cola: Compute-efficient pre-training of llms via low-rank activation
    Ziyue Liu*, Ruijie Zhang*, Zhengyang Wang*, and 5 more authors
    In EMNLP 2025 (oral), 2025
  2. NeurIPS
    LaX: Boosting Low-Rank Training of Foundation Models via Latent Crossing
    Ruijie Zhang, Ziyue Liu, Zhengyang Wang, and 1 more author
    In NeurIPS 2025, 2025
  3. MLSys
    BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
    Zhengyang Wang*, Ziyue Liu*, Ruijie Zhang, and 5 more authors
    In MLSys 2026, 2025
  4. MLSys
    SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
    Jiayi Tian, Seyedarmin Azizi, Yequan Zhao, and 7 more authors
    In MLSys 2026, 2025

2023

  1. MICRO
    Rm-stc: Row-merge dataflow inspired gpu sparse tensor core for energy-efficient sparse acceleration
    Guyue Huang, Zhengyang Wang, Po-An Tsai, and 3 more authors
    In MICRO 2023. Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, Oct 2023