Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents
A lightweight memory construction mechanism with dynamic evolution, giving LLM agents persistent and continually refined memory at low overhead.
I am an Associate Professor at the Research Center for Cognitive Intelligence Systems and the State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (ICT, CAS), where I also serve as a master’s supervisor. From August 2025 to August 2026 I was a visiting scholar at the NExT++ Research Centre, National University of Singapore, hosted by Prof. Tat-Seng Chua. I received my Ph.D. from ICT, CAS in 2021 under the supervision of Prof. Xueqi Cheng and Prof. Huawei Shen, and received my B.S. from Harbin Institute of Technology (Weihai) in 2016. My earlier work centers on graph representation learning, in particular spectral graph neural networks, with applications in financial risk control and social computing. My current research focuses on large language models and LLM agents — inference-time alignment, process reward modeling, memory and planning for agents, and multi-agent social simulation.
We are looking for students (third/fourth year undergraduates or master's students) as interns to collaborate with us in research, as well as assistant professors and postdocs working on large language models and LLM agents. If you are interested, please feel free to contact me at any time.
A lightweight memory construction mechanism with dynamic evolution, giving LLM agents persistent and continually refined memory at low overhead.
Negative reasoning samples — the wrong chains models usually discard — turn out to carry the signal that drives out-of-domain generalization.
A hierarchical demographic tree generates agent populations that adapt to the topic under discussion, making social simulation both faithful and controllable.
An omni-multimodal assistant that understands audio-visual streams as they arrive, so it can respond mid-stream instead of waiting for the input to finish.
Query-level workflow generation costs far more than task-level yet barely helps — a single well-optimized workflow already covers most queries.
We show that a weak teacher — far smaller and less performant than the student — can substantially boost student reasoning, without expert supervision or costly RL.
Multiple personality traits are composed at decoding time, letting one LLM express a target persona without any personality-specific training.
We reformulate alignment as an iterative optimization over an energy function on logits in continuous space, defined by the optimal RLHF policy — achieving deep alignment without retraining.
A unified framework that treats instruction-based image generation and editing as one task, letting the two settings reinforce each other instead of being trained apart.