个人介绍 / About

About Me

I am now a PhD candidate in Computer Science Department at The Hong Kong University of Science and Technology (HKUST), advised by Prof. Xiaofang Zhou and Prof. Long Chen. My current research focuses on enhancing 3D spatial intelligence in MLLMs. My specific interest lies in improving performance on challenging spatial reasoning tasks for current Video MLLM/VLMs. My broader interests include multimodal representation learning, large data systems, and embodied AI in 3D environments.

News

2026.9.26 - 🎉🎉Our work《SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs》is accepted by NeurIPS2026!🎉🎉
2026.8.21 - Grats for《LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention》 being accepted by EMNLP2026!
2026.4.7 - Grats for《ReTRE: Benchmarking LLM Transfer Robustness with Structure-Preserving Variants》 being accepted by ACL2026!
2025.9.26 - Passed PhD Qualifying Exam, I’m now a PhD Candidate!🎉
2025.5.20 - End of semaster, grateful for Machine Learning course and Machine Learning for 3D course where I both received (>=A) scores.
2024.12.20 - 2024秋季学期结束, 所有课程都高分通过, 课程压力大大减少了.
2024.9.4 - 2024秋季学期开始,DSF-Spatiotemporal Analytics创建。
2024.7.2 - 🎉Our work《High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse Rendering》is accepted by ECCV24!🎉
2024.6.5 - Since 2005,HKUSTCSEDB seminar officially restored!
2024.5.30 - 2024春季学期课程终于结束了!!
2024.5.17 - ECCV24 Rebuttal finished!!Beginning of the Finals!!

Publications

  • SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs. NeurIPS 2026

  • LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention. EMNLP 2026.

  • ReTRE: Benchmarking LLM Transfer Robustness with Structure-Preserving Variants. ACL 2026.

  • Ming, X., Li, J., Ling, J., Zhang, L., Xu, F. (2025). High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse Rendering. In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G. (eds) Computer Vision – ECCV 2024. ECCV 2024. Lecture Notes in Computer Science, vol 15128. Springer, Cham. https://doi.org/10.1007/978-3-031-72897-6_7

  • Li, J., Lyu, L., Shi, J., Zhao, J., Xu, J., Gao, J., He, R., & Sun, Z. (2022). Generating community road network from GPS trajectories via style transfer. Proceedings of the 30th International Conference on Advances in Geographic Information Systems. https://doi.org/10.1145/3557915.3560958

  • A Semantic Segmentation based POI Coordinates Generating Framework for On-demand Food Delivery Service. In Proceedings of the 29th International Conference on Advances in Geographic Information Systems, pp. 379-388. 2021.

  • Unclonable Photonic Crystal Hydrogels with Controllable Encoding Capacity for Anti-counterfeiting, ACS Applied Materials & Interfaces 2021, 14, 2369-2380.

  • Spatial Technology Assessment of Green Space Exposure and Myopia. Ophthalmology. 129. 10.1016/j.ophtha.2021. 07.031.

  • Factors influencing subspecialty choice among medical students: a systematic review and meta-analysis, BMJ Open Mar 2019, 9 (3) e022097

  • Acceptance Sampling Plans with Type-I Hybrid Censoring Scheme of Weibull Distribution. Advances in Applied Mathematics.03.184-191.10. 12677/AAM.2014.34027.

Patent

[1]An Unsupervised Method for Trajectory Generating Road Network Based on Style Transfer

[2]A Route Mining Algorithm Based on Track Points

[3]A Method for Enhancing Natural Language Feature Extraction Model using Knowledge Graph

From the journal

近期文章

查看全部文章
  1. 自我认知提升:啥是“辉格史观”?

    摘要: 简单讨论下辉格史观

  2. 当几十GB的JSON文件挡在面前:我的数据探索工具迭代之旅

    摘要: 一个基于Python和Rust的快速查看超大json数据的科学工具. 工具地址: https://github.com/adrianJW421/JustNiceTools/tree/main/JsonPeeker

  3. 当3D高斯溅射学会“边走边看”:聊聊On-the-Fly GS背后的巧思与未来

    摘要: 笔者最近读到一篇非常有意思的论文,名为《Gaussian On-the-Fly Splatting: A Progressive Framework for Robust Near Real-Time 3DGS Optimization》(论文ID: arXiv:2503.13086)。这篇…

  4. 从一篇论文聊到AI的未来:为什么大模型需要“专家外援”?一次关于SpatialBot的深度思考之旅

    摘要: 基于《SpatialBot: Precise Spatial Understanding with Vision Language Models》引起的的LLM设计哲学探讨

较早的文章