Wuhan Institute of Technology
Lecturer and Master's Supervisor, School of Computer Science and Engineering.
I am a lecturer in the School of Computer Science and Engineering at Wuhan Institute of Technology. I received my Ph.D. in Computer Science and Technology from the University of Science and Technology of China in December 2023, advised by Prof. Zhangjin Huang and Prof. Naijie Gu.
My research focuses on 3D vision and robot perception, with particular interests in category-level 6D object pose estimation, point cloud understanding, human-centered vision, and multimodal intelligence.
Our paper “Rotation-Equivariant and Spatially Factorized Keypoint Modeling for Category-Level Object Pose Estimation” was accepted by ACM Multimedia 2026 (CCF A).
Presented DecomPose at the 43rd International Conference on Machine Learning, ICML 2026 (CCF A), in Seoul.
Our paper DecomPose was accepted by the 43rd International Conference on Machine Learning (ICML 2026) (CCF A).
Following its acceptance by ICML 2026, DecomPose was made available on arXiv.
Our paper on wavelet-guided geometric feature enhancement and multimodal fusion was accepted by Information Fusion, a high-impact JCR Q1 journal.
INKL-Pose, our work on instance-adaptive keypoint learning with local-to-global geometric aggregation, was made available on arXiv.

ACM MM 2026CCF A
Proceedings of the 34th ACM International Conference on Multimedia, 2026.
Paper coming soon

ICML 2026CCF A
Proceedings of the 43rd International Conference on Machine Learning, 2026.

Information FusionJCR Q1High-Impact Journal
Information Fusion, 133: 104344, 2026.

arXiv 2025Preprint
arXiv preprint arXiv:2504.15134, 2025.

Pattern RecognitionCCF B
Pattern Recognition, 145: 109896, 2024.

IEEE TCSVTCCF B
IEEE Transactions on Circuits and Systems for Video Technology, 34(4): 2385–2398, 2024.

Neural NetworksCCF B
Neural Networks, 166: 609–621, 2023.

Object pose estimation, point cloud understanding, geometric representation learning, and robot perception.
Human pose estimation, multi-person motion forecasting, activity understanding, and interaction modeling.
Vision-language representation, semantic guidance, multimodal fusion, and embodied perception.
Lightweight visual inspection, anomaly detection, weak supervision, and deployable vision systems.
I teach courses including Algorithm Design and Analysis, Data Structures Project, and Mobile Internet. I welcome students interested in 3D computer vision, human-centered vision, and multimodal perception.
Lecturer and Master's Supervisor, School of Computer Science and Engineering.
Ph.D. in Computer Science and Technology. Advisors: Prof. Zhangjin Huang and Prof. Naijie Gu.
B.Eng. in Internet of Things.