Publications

2026

  1. EMNLP
    m3ec_emnlp26.png
    M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Event Coreference Resolution
    Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Aafiya Shamshad Hussain, Hani Alomari, Alvi Md Ishmam, Chia-Wei Tang, Ali Asgarov, and Chris Thomas
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  2. EMNLP
    nest_emnlp26.png
    NEST: Narrative Event Structures in Time for Long Video Understanding
    Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker, Hani Alomari, Chia-Wei Tang, Anushka Sivakumar, Zaber Ibn Abdul Hakim, Shaurya Mallampati, and Chris Thomas
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  3. EMNLP
    lengthbias_emnlp26.png
    Investigating Length Bias and Robustness in LVLMs for Multiple-Choice Question Answering
    Md. Atabuzzaman, Hani Alomari, and Chris Thomas
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  4. EMNLP
    diffusion_emnlp26.png
    Reliability Challenges in Diffusion Vision–Language Models
    Md. Atabuzzaman and Chris Thomas
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  5. EMNLP
    introconformal_emnlp26.png
    IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals
    Md. Atabuzzaman, Christian Alexander, and Chris Thomas
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  6. CVPR
    immunizing_cvpr26.png
    Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics
    Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang, Alvi Md Ishmam, and Chris Thomas
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  7. CVPR
    lenses_cvpr26.png
    Lenses: Toward Polysemous Vision-Language Understanding
    Hani Alomari, Ali Asgarov, and Chris Thomas
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  8. ECCV
    grounding_eccv26.png
    Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
    Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Chia-Wei Tang, Anuj Karpatne, and Chris Thomas
    In Proceedings of the European Conference on Computer Vision (ECCV), 2026
  9. ACL
    soundbreak_acl26.png
    SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
    Aafiya Hussain, Gaurav Srivastava, Alvi Ishmam, Zaber Hakim, and Chris Thomas
    In Proceedings of the Association for Computational Linguistics (ACL), 2026
  10. CHI
    sensemaking_chi26.png
    Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
    Puqi Zhou, Ali Asgarov, Aafiya Hussain, Wonjoon Park, Amit Paudyal, Sameep Shrestha, Chia-Wei Tang, Michael Lighthiser, Michael Hieb, Xuesu Xiao, Chris Thomas, and Sungsoo Ray Hong
    In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026
  11. AAAI
    lamp_aaai26.png
    LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
    Alvi Md Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, and Chris Thomas
    In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-26), 2026

2025

  1. EMNLP
    ddot_emnlp25.png
    Flexible-length Text Infilling for Discrete Diffusion Models
    Andrew Zhang, Anushka Sivakumar, Chia-Wei Tang, and Chris Thomas
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025
  2. EMNLP
    mcqa_emnlp25.png
    Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
    Md. Atabuzzaman, Ali Asgarov, and Chris Thomas
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025
  3. EMNLP
    Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
    Md. Atabuzzaman, Andrew Zhang, and Chris Thomas
    In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025
  4. EMNLP
    SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
    Anushka Sivakumar, Andrew Zhang, Zaber Ibn Abdul Hakim, and Chris Thomas
    In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025
  5. ACL
    acl25.png
    Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval
    Hani Alomari, Anushka Sivakumar, Andrew Zhang, and Chris Thomas
    In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025
  6. CVPRW
    surgical_cvprw25.png
    Real-Time Ultra-Fine-Grained Surgical Instrument Classification
    Md. Atabuzzaman, Gino DiMatteo, Hani Alomari, Chia-Wei Tang, Connor Hale, Adam E. Goode, David Ryan King, and Chris Thomas
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Fine-Grained Visual Categorization (FGVC) Workshop, 2025
  7. NeurIPSW
    trapping_neuripsw25.png
    Model Immunization by Trapping Harmful Finetuning
    Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Alvi Md Ishmam, Chia-Wei Tang, and Chris Thomas
    In NeurIPS 2025 Workshop on Lock-LLM: Prevent Unauthorized Knowledge Use from LLMs, 2025
  8. arXiv
    PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
    Kazi Hasan Ibn Arif, Sajib Acharjee Dip, Khizar Hussain, Lang Zhang, and Chris Thomas
    arXiv preprint arXiv:2501.12206, 2025
  9. WACV
    Advancing chart question answering with robust chart component recognition
    Hanwen Zheng, Sijia Wang, Chris Thomas, and Lifu Huang
    In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025

2024

  1. NeurIPSW
    ENTER: Event Based Interpretable Reasoning for VideoQA
    Hammad Ayyubi, Junzhang Liu, Ali Asgarov, Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Zhecan Wang, Chia-Wei Tang, Hani Alomari, Md Atabuzzaman, Xudong Lin, Naveen Reddy Dyava, Shih-Fu Chang, and Chris Thomas
    In Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2024, 2024
  2. EMNLP
    emnlp24.png
    M3D: MultiModal MultiDocument Fine-Grained Inconsistency Detection
    Chia-Wei Tang, Ting-Chih Chen, Kiet Nguyen, Kazi Sajeed Mehrab, Alvi Ishmam, and Chris Thomas
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024
  3. ACM
    Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
    Junzhang Liu, Zhecan Wang, Hammad Ayyubi, Haoxuan You, Chris Thomas, Rui Sun, Shih-Fu Chang, and Kai-Wei Chang
    In Proceedings of the 32nd ACM International Conference on Multimedia, 2024
  4. NeurIPS
    neurips24.png
    Journeybench: A challenging one-stop vision-language understanding benchmark of generated images
    Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun, Wenhao Li, Md. Atabuzzaman, Hammad Ayyubi, Haoxuan You, Alvi Md Ishmam, Kai-Wei Chang, Shih-Fu Chang, and Chris Thomas
    In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024
  5. ACL
    acl24.png
    MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking
    Ting-Chih Chen, Chia-Wei Tang, and Christopher Thomas
    In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024
  6. CVPR
    cvpr24.png
    Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
    Alvi Md Ishmam and Christopher Thomas
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
  7. aaai24.png
    Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities
    Hammad Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh, Long Chen, Yulei Niu, Xudong Lin, Xuande Feng, Jaywon Koo, Sounak Ray, and others
    In Proceedings of the AAAI Conference on Artificial Intelligence, 2024

2023

  1. ACL
    Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs
    Mingyang Zhou, Yi Fung, Long Chen, Christopher Thomas, Heng Ji, and Shih-Fu Chang
    In Findings of the Association for Computational Linguistics: ACL 2023, 2023

2022

  1. ECCV
    fgve.png
    Fine-Grained Visual Entailment
    Christopher Thomas, Yipeng Zhang, and Shih-Fu Chang
    In Proceedings of the European Conference on Computer Vision, 2022
  2. Community implications for gun violence prevention during co-occurring pandemics; a qualitative and computational analysis study
    Desmond U. Patton, Nathan Aguilar, Aviv Y. Landau, Chris Thomas, Rachel Kagan, Tianai Ren, Eric Stoneberg, Timothy Wang, Daniel Halmos, Anish Saha, Amith Ananthram, and Kathleen McKeown
    Preventive Medicine, 2022
  3. CVPRW
    complementarity.png
    Emphasizing Complementary Samples for Non-Literal Cross-Modal Retrieval
    Christopher Thomas and Adriana Kovashka
    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2022
  4. TPAMI
    Learning to Overcome Noise in Weak Caption Supervision for Object Detection
    Mesut Erhan Unal, Keren Ye, Mingda Zhang, Christopher Thomas, Adriana Kovashka, Wei Li, Danfeng Qin, and Jesse Berent
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022
  5. EMNLP
    Weakly-Supervised Temporal Article Grounding
    Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, and Shih-Fu Chang
    In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP2022), 2022

2021

  1. ACL
    infosurgeon.png
    InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection
    Yi Fung, Christopher Thomas, Revanth Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, and Avi Sil
    In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021
  2. IJCV
    Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task
    Christopher Thomas and Adriana Kovashka
    International Journal of Computer Vision, 2021
  3. EMNLP
    Joint Multimedia Event Extraction from Video and Article
    Brian Chen, Xudong Lin, Christopher Thomas, Manling Li, Shoya Yoshida, Lovish Chum, Heng Ji, and Shih-Fu Chang
    In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP2021) Findings, 2021

2020

  1. ECCV
    preserving_semantic_neighborhoods.png
    Preserving Semantic Neighborhoods for Robust Cross-modal Retrieval
    Christopher Thomas and Adriana Kovashka
    In Proceedings of the European Conference on Computer Vision (ECCV), 2020
  2. Modeling Visual Rhetoric and Semantics in Multimedia
    Christopher Thomas
    University of Pittsburgh, 2020
  3. arXiv
    Learning to Transfer Visual Effects from Videos to Images
    Christopher Thomas, Yale Song, and Adriana Kovashka
    arXiv preprint arXiv:2012.01642, 2020

2019

  1. NeurIPS
    politics_home_img3.png
    Predicting the politics of an image using webly supervised data
    Christopher Thomas and Adriana Kovashka
    In Advances in Neural Information Processing Systems (NeurIPS 2019), 2019

2018

  1. BMVC
    faces.png
    Persuasive faces: generating faces in advertisements
    Christopher Thomas and Adriana Kovashka
    In Proceedings of the British Machine Vision Conference, 2018
  2. ACCV
    Artistic object recognition by unsupervised style adaptation
    Christopher Thomas and Adriana Kovashka
    In Asian Conference on Computer Vision, 2018

2017

  1. CVPR
    ads_dataset_concept.png
    Automatic understanding of image and video advertisements
    Zaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, and Adriana Kovashka
    In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017

2016

  1. CVPR
    photo_project_hine.jpg
    Seeing Behind the Camera: Identifying the Authorship of a Photograph
    Christopher Thomas and Adriana Kovashka
    In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016
  2. arXiv
    Opensalicon: An open source implementation of the salicon saliency model
    Christopher Thomas
    arXiv preprint arXiv:1606.00110, 2016
  3. CVPRW
    A Visual Attention Algorithm Designed for Coupled Oscillator Acceleration
    Christopher Thomas, Adriana Kovashka, Donald Chiarulli, and Steven Levitan
    In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016

2015

  1. arXiv
    Hand Posture’s Effect on Touch Screen Text Input Behaviors: A Touch Area Based Study
    Christopher Thomas and Brandon Jennings
    arXiv preprint arXiv:1504.02134, 2015
  2. SEKE
    Application of Slow Intelligence Framework for Smart Pet Care System Design
    Shi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, and Christopher Lee Thomas
    In Software Engineering and Knowledge Engineering (SEKE 2015), 2015

2014

  1. INLG
    TBI-Doc: Generating patient & clinician reports from brain imaging data
    Pamela Jordan, Nancy Green, Christopher Thomas, and Susan Holm
    In Proceedings of the 8th International Natural Language Generation Conference (INLG), 2014
  2. Student Response Analysis
    Sean Myers, Timothy Parenti, and Chris Thomas
    2014