Ph.D. Candidate, Artificial Intelligence
The Hong Kong University of Science and Technology (Guangzhou)
Ph.D.Ph.D. Candidate in Artificial Intelligence, HKUST(GZ)
EducationB.E. in Automation, Tsinghua University
Hi! I am a Ph.D. Candidate in Artificial Intelligence at the Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Hui Xiong.
I received my B.E. in Automation from Tsinghua University in 2022. My research combines reinforcement learning, multimodal behavior learning, robot control, and embodied intelligence.
Academic training in artificial intelligence and automation.
The Hong Kong University of Science and Technology (Guangzhou)
Tsinghua University
Learning agents that acquire knowledge and skills through physical interaction.
I aim to build embodied agents that explore autonomously, discover reusable physical and causal structure, and improve continuously through interaction.
Reusable skills and informative interactions without dense supervision.
Efficient policies whose behavior can be selected, adapted, and improved.
Action-conditioned world models that accumulate knowledge without forgetting.
I am exploring embodied scientific discovery: robots autonomously conduct physical experiments, recover relationships among force, mass, acceleration, friction, and contact, then use them for prediction and control.
Interests: Reinforcement Learning · Robot Learning · Multimodal Imitation · World Models · Continual Learning · Embodied Intelligence
Selected first-author and collaborative research. My name is shown in bold.
Source-intervenable flow policies for controllable multimodal robot behavior.
SOMA equips VLA policies with persistent spatial-semantic memory for manipulating targets beyond the current camera view.
Q-guided priors and explicit entropy control for efficient continuous-control policies.
NeuroVLA organizes semantic planning, adaptive motor control, and fast reflexes into a cortex-cerebellum-spinal hierarchy.
A regret-aware min-max framework for improving skill-learning efficiency and diversity.
Multi-turn recommendation that learns latent user preferences through reinforcement learning.
Selected open-source work in embodied intelligence.
Core contributor to an open-source embodied-intelligence framework covering VLA architectures, world models, continual learning, reinforcement learning, shared training, and evaluation infrastructure.
Research and engineering experience, ordered from most recent.
AI² Robotics · Shenzhen
Vision-language-action models, reinforcement learning, humanoid deployment, and open-source research.
Shanghai AI Laboratory · Shanghai
Continual learning, robotic VLA deployment, and unsupervised skill discovery.
OPPO · Shenzhen
Multimodal interactive recommendation with reinforcement learning.
Microsoft Research Asia · Remote
Machine learning, decentralized finance security, and blockchain applications.
Xiaomi · Beijing
Deep click-through-rate prediction and advertising algorithms.
Selected recognition.