Photo of Runhao Zeng

Research Group

Runhao Zeng 曾润浩

Associate Professor, Artificial Intelligence Research Institute
Shenzhen MSU-BIT University

Our group studies embodied affective intelligence — enabling AI systems to understand what people do, infer how they feel and why, and actively perceive and interact with humans in dynamic environments. Our research builds from human action understanding to affective understanding and ultimately to embodied affective intelligence.

I received my Bachelor’s Degree in Automation in 2015 and my Ph.D. in Software Engineering in 2021, both from the South China University of Technology, under the supervision of Prof. Mingkui Tan.

Human Action → Affect → Embodiment
01 / FOUNDATION

Human Action Understanding

Our earlier and continuing work asks how machines can reliably understand what people do over time. We study temporal action localization, video grounding, long-video understanding, and robust and efficient video models — providing the perceptual foundation for reasoning about human affect.

More Work

Selected work across temporal localization, grounding, long video, robustness, and efficiency.

02 / CORE

Affective Understanding

Behavior is observable; affect is latent. We study how machines can move beyond recognizing visible actions to infer affective states from body behavior, temporal context, multimodal evidence, and the events behind observable expressions — including cases where appearance alone can be misleading.

More Work

Additional work on affective behavior, multimodal affect, and psychological-state understanding.

2026 PR
Towards Stable Cross-Domain Depression Recognition under Missing Modalities Multimodal Affect
2025 IoTJ
Skeleton-Based Pre-Training with Discrete Labels for Emotion Recognition in IoT Environments Body Affect
03 / CURRENT DIRECTION

Embodied Affective Intelligence

Real-world affect is not always visible from a fixed viewpoint. We therefore study agents that can actively observe, move, seek evidence, and interact to reduce uncertainty about human affect. This direction connects affective computing with embodied perception, simulation, robot learning, and interactive intelligence.

Current focus. Moving affective computing from passive recognition on fixed datasets toward controllable, interactive, and reproducible embodied environments.

More Work

Additional work in embodied perception, navigation, and robot learning.

2026 AAAI
UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model VLN
2026 AAAI
Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators Robot Learning
2025 TMM
Source-Free Elastic Model Adaptation for Vision-and-Language Navigation VLN
2022 NeurIPS
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation VLN

Selected Applications

AI for Health and Human-Centered Applications

Beyond our main research trajectory, we are interested in applying multimodal perception and learning to selected healthcare and human-centered problems.

MemSAM
CVPR 2024 · Best Paper Award Finalist

MemSAM: Echocardiography Video Segmentation

Adapting foundation segmentation models to noisy, temporally evolving cardiac ultrasound video.

Paper
Flexible triboelectric sensor array
Nano Energy 2024

Flexible Sensing for Human–Machine Interaction

Combining flexible sensing with intelligent recognition for human motion and interaction.

Paper

Publications

Selected Publications

Publications are organized around the research trajectory of human action → affect → embodiment, with additional work in multimodal AI and human-centered applications.

2026
Pattern Recognition

Amplitude Exchanging Network for Unsupervised Underwater Image EnhancementVision

Runhao Zeng, Xionglin Zhu, Wenfu Peng, Jiezhang Cao, Yong Guo, Zhihua Wang, Qiuping Jiang

2026
arXiv

Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language ModelsMLLM

Jingrui Zhang, Feng Liang, Yong Zhang, Wei Wang, Runhao Zeng, Xiping Hu

2025
TCSVT

Toward Long Video Understanding via Fine-Detailed Video Story GenerationLong Video

Zeng You, Zhiquan Wen, Yaofo Chen, Xin Li, Runhao Zeng, Yaowei Wang, Mingkui Tan

2025
AAAI

Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training LayersTTA

Qi Deng, Shuaicheng Niu, Ronghao Zhang, Yaofo Chen, Runhao Zeng, Jian Chen, Xiping Hu

2025
TMM

Source-Free Elastic Model Adaptation for Vision-and-Language NavigationVLN

Mingkui Tan, Peihao Chen, Hongyan Zhi, Jiajie Mai, Benjamin Rosman, Dongyu Ji, Runhao Zeng

2022
NeurIPS

Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationVLN

Peihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng, Thomas Li, Mingkui Tan, Chuang Gan

2020
ACM MM

Cross-modal Relation-aware Networks for Audio-visual Event LocalizationMultimodal Video

Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, Chuang Gan

2019
TMM

Relation Attention for Temporal Action LocalizationTAD

Peihao Chen, Chuang Gan, Guangyao Shen, Wenbing Huang, Runhao Zeng, Mingkui Tan

This page highlights publications most relevant to the research story above. For the complete and continuously updated record, please see Google Scholar or DBLP.

Updates

Recent News

  • 2026/07 Our extended work on the temporal robustness of action localization was accepted by IJCV.
  • 2026/03 Two papers on intelligent healthcare (videofluoroscopy) and keypoint detection were accepted by ICME 2026.
  • 2026/02 Our multimodal large-model work on depression recognition was accepted by Pattern Recognition.
  • 2026/02 Our work on test-time adaptation for generalized temporal action localization was accepted by IEEE TCSVT.
  • 2026/02 Our research received the Wu Wenjun AI Science & Technology Award, First Prize (Natural Science).
  • 2026/01 Our work on unsupervised underwater image enhancement was accepted by Pattern Recognition.
  • 2026/01 Our research received the Second Prize of the Ministry of Education Natural Science Award.
  • 2026/01 One paper on video retrieval was accepted by ICASSP 2026, and one paper on zero-shot group activity recognition was accepted by Pattern Recognition.
  • 2025/11 Three papers on video understanding and embodied intelligence were accepted by AAAI 2026.
  • 2025/06 Our work OVG-HQ was accepted by ICCV 2025.