Chris Thomas
Assistant Professor, Department of Computer Science, Virginia Tech
COMPUTER VISION · MULTIMODAL AI · NATURAL LANGUAGE PROCESSING
I am an Assistant Professor in the Department of Computer Science at Virginia Tech. My research lies at the intersection of computer vision, multimodal AI, machine learning, and natural language processing. I focus on building trustworthy AI systems that can reason reliably over large, heterogeneous collections of images, video, text, and audio. A theme that unifies my recent work is the use of structure, both in how multimodal content is represented and in how models are controlled during training and inference. I use this structure to improve reliability, fine-grained understanding, and resilience to failure. Recent projects include defending vision-language models against adversarial attacks and harmful fine-tuning, extracting events and narratives from large multimodal collections, and embodied perception for robots operating with degraded sensing. I am affiliated with the Sanghani Center for Artificial Intelligence and Data Analytics.
Prior to joining Virginia Tech, I was a postdoctoral researcher at Columbia University working with Professor Shih-Fu Chang. My Ph.D. advisor was Professor Adriana Kovashka.
Note: I am currently recruiting strong and motivated students to join our growing group! Please visit the Prospective Students page for details.
News
Earlier news- AUG 2026 Five papers accepted to the EMNLP 2026 main conference. The papers span multimodal event understanding, long video understanding, and the reliability and factuality of vision-language models.
- JUL 2026 My Ph.D. student Hani Alomari has been awarded an Amazon Fellowship. Congratulations to Hani!
- JUN 2026 Named an Outstanding Area Chair at CVPR 2026. I also gave an invited talk at the CVPR 2026 Area Chair workshop.
- JUN 2026 Our paper SoundBreak was accepted to ACL 2026. The paper shows that adversarial changes to only the audio channel can cause severe failures in models that jointly process audio, video, and text.
- JUN 2026 Our paper on reasoning-guided part-level visual grounding was accepted to ECCV 2026. The paper uses reinforcement learning to teach multimodal models to locate fine-grained object parts by first reasoning about the whole object.
- JUN 2026 Our paper on multi-robot ground video sensemaking was accepted to CHI 2026. Working with public safety professionals, we designed and evaluated tools that help operators make sense of video from fleets of ground robots.
- MAR 2026 Three of my Ph.D. students secured summer research internships at Amazon, Adobe, and Futurewei. Very proud of their hard work!
- FEB 2026 Two papers accepted to CVPR 2026. One paper introduces a new approach for understanding and retrieving images with multiple meanings. The other develops new techniques to prevent misuse of models for harmful tasks.
Research
Trustworthy and secure multimodal AI
Adversarial robustness and red teaming for multimodal models, post-training defenses and model immunization, and inference that is grounded, calibrated, and controllable.
Semantic structure for multimodal understanding
Multimedia information extraction across documents and modalities, event and narrative structure, and benchmarks that expose failure modes in multimodal reasoning.
Embodied AI and real-world systems
Multi-robot video sensemaking with public safety professionals and audio-visual perception for embodied systems operating with degraded or unreliable sensors.
Recent publications
All publications
Students in the group work on computer vision, multimodal AI, and natural language processing, and present their work at the main conferences in these areas. Former members of the group have gone on to positions at Amazon, Apple, Argonne National Laboratory, Capital One, Microsoft, and TikTok.
Teaching
| CS 6804 | World Models | Fall 2026 |
| CS 5814 | Introduction to Deep Learning | Spring 2026, Spring 2024 |
| CS 5864 | Learning-based Computer Vision | Fall 2025, Fall 2023 |
| CS 6804 | Multimodal Vision | Fall 2024, Spring 2023 |
| CS 4894 | Introduction to Deep Learning and Computer Vision | Spring 2025 |