Zhi-Qi Cheng, Ph.D.
About
Zhi-Qi Cheng is an Assistant Professor of Computer Science & Systems in the School of Engineering & Technology at the University of Washington Tacoma. He directs the Multimodal Intelligence Lab (MILab) and is a UW Graduate Faculty member with doctoral endorsement.
Research
Multimodal AI · World Models · Intelligent Transportation
He develops multimodal AI and world models that combine 3D geometry, semantic understanding, and generative modeling for perception, spatial reasoning, prediction, and planning from incomplete observations. Applications include visual navigation in dynamic environments and transportation infrastructure mapping, monitoring, and management.
Two projects funded by the U.S. Department of Transportation through PacTrans support his research in 3D urban scene reconstruction and transportation asset mapping.
Background
Before joining UW, he spent seven years at Carnegie Mellon University’s Language Technologies Institute, School of Computer Science, as a Research Associate (2017–2019), Postdoctoral Research Associate (2019–2022), and Project Scientist (2022–2024). His doctoral and postdoctoral research advisors were Professor Alexander G. Hauptmann and Professor Teruko Mitamura, respectively.
He was a core technical lead for CMU’s DARPA KAIROS and KAIROS Plus projects (2019–2024) and contributed to DARPA AIDA (2018–2023), DARPA GAILA (2019–2022), IARPA DIVA (2017–2021), and NIST PSIAP (2017–2019).
His industry experience includes Visiting Faculty Researcher at Meta AI (2025) and Research Intern at Microsoft Research (2019) and Google Brain (2018). He also contributed technical analysis to The Washington Post’s reporting recognized with the 2022 Pulitzer Prize for Public Service.
Student advising & opportunities
He primarily recruits Ph.D. and master’s students through UW’s Computer Science & Systems (CSS) programs. He can also co-advise students in CSE, ECE, and related UW programs and serve on their doctoral supervisory committees.
He welcomes prospective and current students interested in multimodal AI, world models, or intelligent transportation. Undergraduate students across UW are also welcome to inquire about research. Email zhiqics@uw.edu with a CV and brief research interests to discuss opportunities and the appropriate graduate program.
CSS Ph.D. program · CSS M.S. program · Student opportunities & achievements
Teaching
His teaching spans graduate and undergraduate courses in machine learning, multimodal AI, algorithms, robotics, and computer graphics.
- TCSS 590: Special Topics — Vision–Language Models
- TCSS 543: Advanced Algorithms
- TCSS 455: Introduction to Machine Learning
- TCSS 437: Mobile Robotics
- TCSS 458: Computer Graphics
Students across UW are welcome to take his courses. See the UW cross-campus enrollment policy. Research supervision includes undergraduate research, graduate independent study, master’s theses and design projects, and doctoral research. See courses and research credit.
Student achievements & outreach
His Ph.D. students Yifei Dong and Fengyi Wu received Graduate School Top Scholar Awards in 2025 and 2026, respectively. Dong also received the 2025 Carwein–Andrews Ph.D. Fellowship.
Dr. Cheng has served as a judge and advisor for Regeneron ISEF and an advisor for the S.-T. Yau High School Science Award. His high-school mentees include Jasmine Liu, recipient of the ACM Second Award and a Fourth Award in Robotics and Intelligent Machines at ISEF 2023, and Yulun (Alan) Wu, recipient of the 2020 Yau Bronze Prize in Computer Science.
His high-school mentees have been admitted to MIT, UChicago, UIUC, Imperial College London, and Harvey Mudd College.
Recognition & service
Selected recognition includes the Intel Ph.D. Fellowship, CSC–IBM Outstanding Student Scholarship, ICCV Outstanding Reviewer, and a CVPR Anti-UAV Workshop Best Paper Award. At UW, he was selected for the 2025–26 Undergraduate Research Faculty Fellows program supported by the Henry Luce Foundation. His academic service includes Area Chair for CVPR 2027 and NAACL 2025. He also serves on the Program Committee of the RoboWorld Challenge 2026.
Academic service & faculty development · Media collaborations
Illustrated selected publications · MILab publication list · Google Scholar
World models & visual navigation
- Language-Conditioned World Modeling for Visual Navigation. NeurIPS 2026 Oral.
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight. ECCV 2026.
- HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions. IROS 2026.
- Human-Aware Vision-and-Language Navigation. NeurIPS 2024 · Datasets and Benchmarks · Spotlight.
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning. Findings of ACL 2026.
Visual perception & retrieval
- Rethinking Spatial Invariance of Convolutional Networks for Object Counting. CVPR 2022.
- Learning Spatial Awareness to Improve Crowd Counting. ICCV 2019 Oral.
- BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition. CVPR 2024.
- HDFormer: High-Order Directed Transformer for 3D Human Pose Estimation. IJCAI 2023.
- DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving. IJCAI 2023.
- Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images. CVPR 2017.
- Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model Search. ACM Multimedia 2018.
- Video eCommerce++: Toward Large Scale Online Video Advertising. IEEE Transactions on Multimedia 2017.
- Video eCommerce: Towards Online Video Advertising. ACM Multimedia 2016 · ACM SCF Best Student Paper Award.
Multimodal reasoning & generation
- ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic Rules. ICCV 2023.
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention Refinement. ACM Multimedia 2022.
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning. NeurIPS 2024.
- MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis. ICLR 2025.
- Combo: Co-Speech Holistic 3D Human Motion Generation and Efficient Customizable Adaptation in Harmony. IEEE TPAMI 2025 · Early access.
- StableAnimator: High-Quality Identity-Preserving Human Image Animation. CVPR 2025.
- SHIELD: LLM-Driven Schema Induction for Predictive Analytics in EV Battery Supply Chain Disruptions. EMNLP 2024 · Industry Track.
- ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding. NAACL 2025.
- Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards. ACL 2026.
- Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions. CVPR 2025 · Anti-UAV Workshop · Best Paper Award.
Efficient & robust learning
- Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding. ICLR 2026 Oral.
- MaxSup: Overcoming Representation Collapse in Label Smoothing. NeurIPS 2025 Oral.
- Towards Calibrated Robust Fine-Tuning of Vision-Language Models. NeurIPS 2024.
- Robust Adaptation of Foundation Models with Black-Box Visual Prompting. IEEE TPAMI 2026.