Xinyu Lian

Xinyu Lian

I am a third-year CS Ph.D. candidate at UIUC, working with Minjia Zhang. I study efficient machine learning systems, with a current focus on long-horizon agentic RL.

My work has been adopted into DeepSpeedGitHub stars and ArcticGitHub stars, and has been used to train production models, including Microsoft’s Phi-3.5-MoE, Snowflake’s Arctic-Text2SQL-R2, and BigScience’s BLOOM-176B.

I am a Microsoft Research Fellow and an Amazon AI PhD Fellow, and my work received the ASPLOS 2026 Best Paper Honorable Mention.

🙋 Feel free to reach out via email if you are interested in my research.


🔥 News

Jun 29, 2026 Excited to open-source Arctic-RL, an RL backend that accelerates training by up to 3.5×.
Apr 1, 2026 Our work SuperOffload received the ASPLOS 2026 Best Paper Honorable Mention.
Mar 2, 2026 I am honored to receive the Microsoft Research Fellowship.
Feb 3, 2025 I am honored to receive the Amazon AI PhD Fellowship.

📃 Selected Papers

  1. EMNLP’26
      Oral  
    Overview figure: Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models
    Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models
    Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing. 2026
  2. ASPLOS’26
    Artifact AvailableArtifact FunctionalArtifact Reusable
    Overview figure: SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
    SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
    Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 2026
    🏆 ASPLOS 2026 Best Paper Honorable Mention
    Featured by PyTorch Official Blog
    Presented at PyTorch Conference 2025, Ray x DeepSpeed Meetup
  3. ATC’25
    Artifact AvailableArtifact FunctionalArtifact Reusable
    Overview figure: Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelism
    Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelism
    Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase,  and Minjia Zhang
    Proceedings of the 2025 USENIX Annual Technical Conference. 2025
    Adopted by Microsoft for Phi-3.5-MoE (42B) Pre-Training
    Adopted by BigScience for BLOOM (176B) Pre-Training
    Adopted by StepFun and Baidu for multimodal model post-training.
    Presented at PyTorch Day 2025, FMS 2025, SDC 2025
  4. ICSE’25
    Artifact AvailableArtifact FunctionalArtifact Reusable
    Overview figure: Large Language Models as Configuration Validators
    Large Language Models as Configuration Validators
    Proceedings of the 47th IEEE/ACM International Conference on Software Engineering. 2025


🎤 Talks

Oct 2025 Talk at PyTorch Conference 2025, San Francisco, on SuperOffload.
Oct 2025 Talk at Ray x DeepSpeed Meetup, San Francisco, on SuperOffload.
Sep 2025 Talk at SNIA Developer Conference (SDC) 2025, Santa Clara, on Universal Checkpointing.
Aug 2025 Talk at Future of Memory and Storage (FMS) 2025, Santa Clara, on Universal Checkpointing.
May 2025 Talk at PyTorch Day France 2025, Paris, on Universal Checkpointing.

🏅 Awards & Honors

Microsoft Research Fellowship ($43K)
Amazon AI PhD Fellowship ($70K)
ASPLOS 2026 Best Paper Honorable Mention
ASPLOS 2026 Travel Grant
USENIX ATC 2025 Travel Grant
SIGSOFT CAPS Travel Grant
UIUC Conference Presentation Award
Outstanding Graduate of Zhejiang Province

🏖️ Service

Organizer: Brett Daniel Software Engineering Seminar
Program Committee: MSR'25
Reviewer: ICML, NeurIPS, TSE
Artifact Evaluation Committee: SOSP'24