Bingbing Xu

About

I am an Associate Professor at the Research Center for Cognitive Intelligence Systems and the State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (ICT, CAS), where I also serve as a master’s supervisor. From August 2025 to August 2026 I was a visiting scholar at the NExT++ Research Centre, National University of Singapore, hosted by Prof. Tat-Seng Chua. I received my Ph.D. from ICT, CAS in 2021 under the supervision of Prof. Xueqi Cheng and Prof. Huawei Shen, and received my B.S. from Harbin Institute of Technology (Weihai) in 2016. My earlier work centers on graph representation learning, in particular spectral graph neural networks, with applications in financial risk control and social computing. My current research focuses on large language models and LLM agents — inference-time alignment, process reward modeling, memory and planning for agents, and multi-agent social simulation.

Openings

We are looking for students (third/fourth year undergraduates or master's students) as interns to collaborate with us in research, as well as assistant professors and postdocs working on large language models and LLM agents. If you are interested, please feel free to contact me at any time.

Selected Publications

All publications →
Illustration of Incentivizing Strong Reasoning from Weak Supervision EACL 2026

Incentivizing Strong Reasoning from Weak Supervision

Yige Yuan, Teng Xiao, Shuchang Tao, Xue Wang, Jinyang Gao, Bolin Ding, Bingbing Xu

We show that a weak teacher — far smaller and less performant than the student — can substantially boost student reasoning, without expert supervision or costly RL.

Illustration of Inference-Time Alignment in Continuous Space NeurIPS 2025

Inference-Time Alignment in Continuous Space

Yige Yuan, Teng Xiao, Yunfan Li, Bingbing Xu, Shuchang Tao, Yunqi Qiu, Huawei Shen, Xueqi Cheng

We reformulate alignment as an iterative optimization over an energy function on logits in continuous space, defined by the optimal RLHF policy — achieving deep alignment without retraining.

Recent News

All news →
  • 2026-08 Completed my one-year visiting scholarship at the NExT++ Centre, National University of Singapore.
  • 2026-05 Five papers accepted at ACL 2026 — three in the Main Conference (Chain-of-Memory, Learning from Mistakes, HAG) and two in Findings (ROMA, Rethinking Agentic Workflow Generation).
  • 2025-12 One paper accepted at WSDM 2026: Multi-Personality Generation of LLMs at Decoding-Time.
  • 2025-09 One paper accepted at NeurIPS 2025: Inference-Time Alignment in Continuous Space.
  • 2025-08 Started a one-year visit at the NExT++ Centre, National University of Singapore, hosted by Prof. Tat-Seng Chua.