ℹ️ Short Bio
Hi, I’m Yuanzhe, a Ph.D. student in Computer and Information Sciences at the Georgia Institute of Technology (Georgia Tech). I am broadly interested in building reliable, adaptive, and efficient AI systems—especially agents that must learn, remember, reason, and act over long horizons. I am open to research collaborations, so please feel free to reach out!
Before joining Georgia Tech, I received my M.S. in Computer Science and Engineering from the University of California, San Diego (UCSD) 🔱 and my B.S. in Artificial Intelligence from Huazhong University of Science and Technology (HUST). I have been fortunate to collaborate with Prof. Yaoqing Yang at Dartmouth College, Prof. Julian McAuley at UC San Diego, and Prof. Zhiting Hu at UC San Diego.
My research spans two complementary directions:
-
Reliable AI agents and long-horizon learning. I develop evaluation harnesses, memory systems, and post-training methods to study how agents acquire, update, retrieve, and use knowledge across extended interactions and complex environments. Recent projects include EarthVerse, MemoryArena (ICML 2026), MemoryAgentBench (ICLR 2026), Mem-$\alpha$, M+ (ICML 2025), K2-Think, and MIRIX.
-
Understanding and improving learning systems. I use mathematical and spectral perspectives to study model structure, training dynamics, generalization, and failure modes, and translate these insights into more efficient training and compression methods. This line includes Spectral Signatures of Large Language Models (KDD 2026), SciML Diagnosis (ICML 2026), FARMS (ICML 2025), and Model Balancing (EMNLP 2024 Oral).
Looking forward, I am particularly interested in agent harnesses and memory for long-horizon learning and decision-making, self-improving systems that learn continually from interaction and feedback, and agent architectures that reliably work with tools, services, and structured data sources such as databases.
🔥 News
-
2026.08: Started my PhD journey.
-
2026.06: I gave a talk about Long-horizon Agent Evalution at Cornell Tech. See Slides
-
2026.05: 🎉🎉🎉 One paper is accepted by KDD 2026. Two papers are accepted by ICML 2026 as Regular! See you at Seoul, South Korea.
-
2026.04: 😁 I graduated from UCSD!
- 2026.01: 🎉🎉 Our paper “Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions” was accepted by ICLR 2026.
- 2025.07: 😁 We open-sourced the MemoryAgentBench. Thanks for the great help from Yu Wang!
- 2025.05: 🎉🎉 Two papers are accepted by ICML 2025 as Poster! See you at Vancouver.
- 2024.09: 🎉🎉 Excited to share that our work “Model Balancing Helps Low-data Training and Fine-tuning” is accepted by EMNLP 2024 as Oral Presentation!
- 2024.06: 😁 I graduated from HUST! 😄 I created my account on OpenReview!
📖 Educations
|
Georgia Institute of Technology (Georgia Tech) Ph.D. in Computer and Information Sciences |
2026.08 - |
|
University of California, San Diego (UCSD) M.S. in Computer Science and Engineering |
2024.09 - 2026.03 |
|
Huazhong University of Science and Technology (HUST) B.S. in Artificial Intelligence, Innovation Experimental Honor Class, Qiming School GPA: 3.91/4.0 |
2020.09 - 2024.06 |
🔧 Industrial Experience
|
Institute of Foundation Models, MBZUAI Research Collaborator, post-training for LLM reasoning. |
2025.06 - 2025.09 |
⚙️ Research Project
# denotes equal contribution
🤔 Reliable AI Agents and Long-Horizon Learning

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
{Yuanzhe Hu#, Yu Wang#}, Julian McAuley
ICLR 2026
Short Summary: MemoryAgentBench is a new benchmark designed to comprehensively evaluate memory agents in LLMs.

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
{Zexue He#, Yu Wang#, Churan Zhi#, Yuanzhe Hu#, Tzu-Ping Chen#, Lang Yin#}, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland
ICML 2026
Short Summary: We present MemoryAreana, a new evaluation gym designed to bridge the gap between isolated recall and execution by benchmarking agents on tasks where memory acquisition and action are tightly coupled.
EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
Preprint
📖 Understanding and Improving Learning Systems

Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Optimization
{Yuxin Wang#, Yuanzhe Hu#, Xiaokun Zhong#, Xiaopeng Wang#}, Haiquan Lu, Tianyu Pang, Michael W. Mahoney, Yujun Yan, Pu Ren, Yaoqing Yang
ICML 2026
Short Summary: A diagnosis framework for Scientific Machine Learning Models.

