He Zhang 张贺 · Zhang He · Zhanghe

Ph.D.Ph.D. Candidate in Artificial Intelligence, HKUST(GZ)

EducationB.E. in Automation, Tsinghua University

Emailzhanghe18@tsinghua.org.cn

Portrait of He Zhang

Hi! I am a Ph.D. Candidate in Artificial Intelligence at the Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Hui Xiong.

I received my B.E. in Automation from Tsinghua University in 2022. My research combines reinforcement learning, multimodal behavior learning, robot control, and embodied intelligence.

Education

Academic training in artificial intelligence and automation.

2022 - Present

Ph.D. Candidate, Artificial Intelligence

The Hong Kong University of Science and Technology (Guangzhou)

2018 - 2022

B.E., Automation

Tsinghua University

Research

Learning agents that acquire knowledge and skills through physical interaction.

I aim to build embodied agents that explore autonomously, discover reusable physical and causal structure, and improve continuously through interaction.

Autonomous Exploration & Skill Discovery

Reusable skills and informative interactions without dense supervision.

Intervenable Multimodal Behavior

Efficient policies whose behavior can be selected, adapted, and improved.

World Models & Lifelong Robot Learning

Action-conditioned world models that accumulate knowledge without forgetting.

Current research plan

I am exploring embodied scientific discovery: robots autonomously conduct physical experiments, recover relationships among force, mass, acceleration, friction, and contact, then use them for prediction and control.

Interests: Reinforcement Learning · Robot Learning · Multimodal Imitation · World Models · Continual Learning · Embodied Intelligence

Publications

Selected first-author and collaborative research. My name is shown in bold.

  1. Under Review2026

    Source-Lifted Flow Matching for Intervenable Multimodal Imitation

    He Zhang, Ying Sun, Hui Xiong

    Source-intervenable flow policies for controllable multimodal robot behavior.

    Source-Lifted Flow Matching method overview
  2. CollaborativearXiv 2026

    Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

    Pengteng Li, Weiyu Guo, He Zhang, Tiefu Cai, Xiao He, Yandong Guo, Hui Xiong

    SOMA equips VLA policies with persistent spatial-semantic memory for manipulating targets beyond the current camera view.

  3. ICLR 20262026

    GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies

    He Zhang, Ying Sun, Hui Xiong

    Q-guided priors and explicit entropy control for efficient continuous-control policies.

    GoldenStart Q-guided prior illustration
  4. CollaborativearXiv 2026

    A Brain-Inspired Embodied Intelligence for Fluid and Fast Reflexive Robotics Control

    Weiyu Guo, He Zhang, Pengteng Li, Tiefu Cai, Ziyang Chen, Yandong Guo, Xiao He, Yongkui Yang, Ying Sun, Hui Xiong

    NeuroVLA organizes semantic planning, adaptive motor control, and fast reflexes into a cortex-cerebellum-spinal hierarchy.

  5. ICML 20252025

    Efficient Skill Discovery via Regret-Aware Optimization

    He Zhang, Ming Zhou, Shaopeng Zhai, Ying Sun, Hui Xiong

    A regret-aware min-max framework for improving skill-learning efficiency and diversity.

    Regret-aware skill discovery overview
  6. ACM MM2023

    Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning

    He Zhang, Ying Sun, Weiyu Guo, Yafei Liu, Haonan Lu, Xiaodong Lin, Hui Xiong

    Multi-turn recommendation that learns latent user preferences through reinforcement learning.

    Interactive interior design recommendation system

Open source

Selected open-source work in embodied intelligence.

AlphaBrain Platform

Core contributor to an open-source embodied-intelligence framework covering VLA architectures, world models, continual learning, reinforcement learning, shared training, and evaluation infrastructure.

  • Developed RLT/RLT_a and PPO/GRPO trainers.
  • Integrated Qwen and π₀.₅ for VLA fine-tuning and online RL.
  • Refactored LIBERO evaluation and distributed execution.
43 public commits · first by commit count View on GitHub
AlphaBrain platform reinforcement learning capabilities

Experience

Research and engineering experience, ordered from most recent.

  1. 2025 — 2026

    Algorithm Engineer

    AI² Robotics · Shenzhen

    Vision-language-action models, reinforcement learning, humanoid deployment, and open-source research.

  2. 2024 — 2025

    Algorithm Intern

    Shanghai AI Laboratory · Shanghai

    Continual learning, robotic VLA deployment, and unsupervised skill discovery.

  3. 2022 — 2023

    Algorithm Intern

    OPPO · Shenzhen

    Multimodal interactive recommendation with reinforcement learning.

  4. 2022 — 2023

    Remote Research Intern

    Microsoft Research Asia · Remote

    Machine learning, decentralized finance security, and blockchain applications.

  5. 2021

    Algorithm Intern

    Xiaomi · Beijing

    Deep click-through-rate prediction and advertising algorithms.

Awards

Selected recognition.