Zhi Zheng’s Homepage
I am a third-year Ph.D. candidate from School of Computing, National University of Singapore, supervised by Prof. Wee Sun Lee, and Prof. Yee Whye Teh from Oxford.
Previously, I finished my undergraduate studies in June 2024 at the Southern University of Science and Technology (SUSTech), China, focusing on Neural Combinatorial Optimization and Reinforcement Learning, supervised by Prof. Zhenkun Wang, Prof. Xin Yao, and Prof. Ke Tang. I am still working with them now.
Research Interests:
My research focuses on the optimization of LLMs and neural networks, where I develop efficient and scalable methods for optimizing neural agents and harness them to solve challenging optimization problems. My research agenda is organized around three core objectives:
Scalable Fine-tuning: Developing scalable fine-tuning methods for neural agents, including Evolutionary Strategies (ES) (e.g., Agentic-ESOpt, Understanding ES, Hyper-ES) and Reinforcement Learning (RL) (e.g., SofT-GRPO).
Efficiency: Improving the efficiency of neural agents by reducing the computational and representational overhead of reasoning and memory. (e.g., ATP-Latent, Latent Memory)
Problem Solving: Leveraging neural agents to tackle challenging optimization problems via agentic test-time compute (e.g., MCTS-AHD, APEX) or learning heuristics. (e.g., UDC, DPN)
I am willing to discuss the above topics via email!
My CV is here: Zhi Zheng’s Curriculum Vitae.
Email: zhi.zheng@u.nus.edu/ Google Scholar Profile / Github / alphaXiv
Selected Publications:
Arxiv Preprints:
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, and Wee Sun Lee;
Arxiv, 2026. [Project Page & Paper & Code].One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA
Zhi Zheng, Ziqiao Meng, Hao Luan, Wei Liu, and Wee Sun Lee;
Arxiv, ICML2026 Workshop @ Efficient Multimodal Question Answering, 2026. [Paper & Code].SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
Zhi Zheng, Yu Gu, Wei Liu, Yee Whye Teh, and Wee Sun Lee;
Arxiv, 2025. [Paper & Code].Beyond Imitation: Reinforcement Learning for Active Latent Planning
Zhi Zheng and Wee Sun Lee;
Arxiv, 2026. [Paper & Code].
Accepted Conference Papers:
Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design
Zhi Zheng, Zhuoliang Xie, Zhenkun Wang, and Bryan Hooi;
International Conference on Machine Learning (ICML), 2025. [Paper & Code].- UDC: A Unified Neural Divide-and-Conquer Framework for Large-Scale Combinatorial Optimization Problems
Zhi Zheng, Changliang Zhou, Xialiang Tong, Mingxuan Yuan, Zhenkun Wang;
Advances in Neural Information Processing Systems (NeurIPS), 2024. [Paper & Code]. DPN: Decoupling Partition and Navigation for Neural Solvers of Min-max Vehicle Routing Problems
Zhi Zheng*, Shunyu Yao*, Zhenkun Wang, Xialiang Tong, Mingxuan Yuan, Ke Tang;
International Conference on Machine Learning (ICML), 2024. [Paper & Code].- Learning Encodings of Constructive Neural Combinatorial Optimization Needs to Regret
Rui sun*, Zhi Zheng*, Zhenkun Wang;
38th AAAI Conference on Artificial Intelligence (AAAI), 2024. [Paper & Code].
Accepted Journal Articles:
- Pareto Improver: Learning Improvement Heuristics for Multi-Objective Route Planning
Zhi Zheng*, Shunyu Yao*, Genghui Li, Linxi Han, and Zhenkun Wang;
IEEE Transactions on Intelligent Transportation Systems (T-ITS), 2023. [Paper & Code].
* for equal contribution
Teaching:
Teaching Assistant for CS3244, School of Computing, NUS (2025 Spring).
Teaching Assistant for CS3263, School of Computing, NUS (2025 Fall).
