Chris Thomas

Assistant Professor, Department of Computer Science, Virginia Tech

COMPUTER VISION · MULTIMODAL AI · NATURAL LANGUAGE PROCESSING

Chris Thomas

I am an Assistant Professor in the Department of Computer Science at Virginia Tech. My research lies at the intersection of computer vision, multimodal AI, machine learning, and natural language processing. I focus on building trustworthy AI systems that can reason reliably over large, heterogeneous collections of images, video, text, and audio. A theme that unifies my recent work is the use of structure, both in how multimodal content is represented and in how models are controlled during training and inference. I use this structure to improve reliability, fine-grained understanding, and resilience to failure. Recent projects include defending vision-language models against adversarial attacks and harmful fine-tuning, extracting events and narratives from large multimodal collections, and embodied perception for robots operating with degraded sensing. I am affiliated with the Sanghani Center for Artificial Intelligence and Data Analytics.

Prior to joining Virginia Tech, I was a postdoctoral researcher at Columbia University working with Professor Shih-Fu Chang. My Ph.D. advisor was Professor Adriana Kovashka.

Note: I am currently recruiting strong and motivated students to join our growing group! Please visit the Prospective Students page for details.

  • AUG 2026 Five papers accepted to the EMNLP 2026 main conference. The papers span multimodal event understanding, long video understanding, and the reliability and factuality of vision-language models.
  • JUL 2026 My Ph.D. student Hani Alomari has been awarded an Amazon Fellowship. Congratulations to Hani!
  • JUN 2026 Named an Outstanding Area Chair at CVPR 2026. I also gave an invited talk at the CVPR 2026 Area Chair workshop.
  • JUN 2026 Our paper SoundBreak was accepted to ACL 2026. The paper shows that adversarial changes to only the audio channel can cause severe failures in models that jointly process audio, video, and text.
  • JUN 2026 Our paper on reasoning-guided part-level visual grounding was accepted to ECCV 2026. The paper uses reinforcement learning to teach multimodal models to locate fine-grained object parts by first reasoning about the whole object.
  • JUN 2026 Our paper on multi-robot ground video sensemaking was accepted to CHI 2026. Working with public safety professionals, we designed and evaluated tools that help operators make sense of video from fleets of ground robots.
  • MAR 2026 Three of my Ph.D. students secured summer research internships at Amazon, Adobe, and Futurewei. Very proud of their hard work!
  • FEB 2026 Two papers accepted to CVPR 2026. One paper introduces a new approach for understanding and retrieving images with multiple meanings. The other develops new techniques to prevent misuse of models for harmful tasks.

Research

Trustworthy and secure multimodal AI

Adversarial robustness and red teaming for multimodal models, post-training defenses and model immunization, and inference that is grounded, calibrated, and controllable.

Semantic structure for multimodal understanding

Multimedia information extraction across documents and modalities, event and narrative structure, and benchmarks that expose failure modes in multimodal reasoning.

Embodied AI and real-world systems

Multi-robot video sensemaking with public safety professionals and audio-visual perception for embodied systems operating with degraded or unreliable sensors.

Recent publications

All publications
  1. EMNLP 2026

    M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Event Coreference Resolution

    Z. Hakim, N. Sarker, A. Hussain, H. Alomari, A. Ishmam, C. Tang, A. Asgarov, Chris Thomas

  2. EMNLP 2026

    NEST: Narrative Event Structures in Time for Long Video Understanding

    A. Asgarov, K. Narasimhan, N. Sarker, H. Alomari, C. Tang, A. Sivakumar, Z. Hakim, S. Mallampati, Chris Thomas

  3. EMNLP 2026

    Investigating Length Bias and Robustness in LVLMs for Multiple-Choice Question Answering

    M. Atabuzzaman, H. Alomari, Chris Thomas

  4. EMNLP 2026

    Reliability Challenges in Diffusion Vision–Language Models

    M. Atabuzzaman, Chris Thomas

  5. EMNLP 2026

    IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals

    M. Atabuzzaman, C. Alexander, Chris Thomas

  6. CVPR 2026Highlight

    Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics

    N. Sarker, Z. Hakim, A. Asgarov, C. Tang, A. Ishmam, Chris Thomas

  7. CVPR 2026

    Lenses: Toward Polysemous Vision-Language Understanding

    H. Alomari, A. Asgarov, Chris Thomas

  8. ECCV 2026

    Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

    K. Mehrab, H. Alomari, N. Sarker, Z. Hakim, C. Tang, A. Karpatne, Chris Thomas

  9. ACL 2026

    SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models

    A. Hussain, G. Srivastava, A. Ishmam, Z. Hakim, Chris Thomas

  10. CHI 2026

    Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals

    P. Zhou, A. Asgarov, A. Hussain, W. Park, A. Paudyal, S. Shrestha, C. Tang, M. Lighthiser, M. Hieb, X. Xiao, Chris Thomas, S. Hong

  11. AAAI 2026

    LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models

    A. Ishmam, N. Sarker, Z. Hakim, Chris Thomas

Md Atabuzzaman
Md Atabuzzaman
Ph.D. student
Hani Alomari
Hani Alomari
Ph.D. student
Ali Asgarov
Ali Asgarov
Ph.D. student
Kazi Mehrab
Kazi Mehrab
Ph.D. student, co-advised with Anuj Karpatne
Mingsi Liao
Mingsi Liao
Ph.D. student, co-advised with Rebecca Cockrum
Members of the group at CVPR 2026 in Denver
MEMBERS OF THE GROUP AT CVPR 2026 IN DENVER

Students in the group work on computer vision, multimodal AI, and natural language processing, and present their work at the main conferences in these areas. Former members of the group have gone on to positions at Amazon, Apple, Argonne National Laboratory, Capital One, Microsoft, and TikTok.

Teaching

CS 6804 World Models Fall 2026
CS 5814 Introduction to Deep Learning Spring 2026, Spring 2024
CS 5864 Learning-based Computer Vision Fall 2025, Fall 2023
CS 6804 Multimodal Vision Fall 2024, Spring 2023
CS 4894 Introduction to Deep Learning and Computer Vision Spring 2025