Publications
2026
- EMNLP
M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Event Coreference ResolutionZaber Ibn Abdul Hakim, Najibul Haque Sarker, Aafiya Shamshad Hussain, Hani Alomari, Alvi Md Ishmam, Chia-Wei Tang, Ali Asgarov, and Chris ThomasIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026Understanding how the same real-world event is represented across text, video, and audio remains a fundamental challenge in information extraction. While prior datasets have explored events within individual modalities or limited modality pairs, a setting that combines all three modalities remains largely unexplored. We introduce M3EC (MultiModal Multidocument Event & Coreference), a new benchmark for multimodal, multidocument event extraction and coreference resolution. It is built from tightly related news clusters that combine eventful videos, audio files, and news articles. We also introduce a semi-automated annotation pipeline, AutoEventAnnotator, to generate automated annotations at scale. Additionally, we construct a human-annotated test set for robust evaluation. To our knowledge, this is the first benchmark to support cluster-level multidocument event extraction and event coreference across text, video, and audio. Evaluation with strong multimodal LLMs shows that the majority of them fail to extract valid events in video and particularly in audio. We further show that models finetuned on the annotations generated by our automated pipeline achieve consistent improvements, including more than +30 F1 on video event extraction, +15 F1 on audio event extraction and +40 CoNLL F1 on end-to-end coreference resolution, highlighting the usefulness of M3EC for training and evaluation in multimodal event understanding.
@inproceedings{hakim2026m3ec, title = {M3EC: A Multimodal Multidocument Benchmark for Event Extraction and Event Coreference Resolution}, author = {Hakim, Zaber Ibn Abdul and Sarker, Najibul Haque and Hussain, Aafiya Shamshad and Alomari, Hani and Ishmam, Alvi Md and Tang, Chia-Wei and Asgarov, Ali and Thomas, Chris}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - EMNLP
NEST: Narrative Event Structures in Time for Long Video UnderstandingAli Asgarov, Kaushik Narasimhan, Najibul Haque Sarker, Hani Alomari, Chia-Wei Tang, Anushka Sivakumar, Zaber Ibn Abdul Hakim, Shaurya Mallampati, and Chris ThomasIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026Recent progress in vision-language models has enabled processing of increasingly long video sequences, but the ability to handle extended token streams does not translate to understanding of narrative structure in long videos. Existing long video benchmarks focus on needle-in-a-haystack retrieval rather than evaluating how low-level actions form events, how events interact across time, and how narratives progress, for example whether a model can connect an early setback such as a job loss to a later relationship breakup, despite long gaps, intervening scenes, or flashbacks that reframe what occurred. We introduce NEST (Narrative Event Structures in Time for Long Video Understanding), a dataset of 1005 full-length movies (avg. 98 minutes), each annotated with 102 multimodal narrative events grounded in visual content, dialogue, and audio. NEST captures multimodal narrative events with structured annotations grounded in visual content, dialogue, and audio, and links them through relations that reflect narrative structure, including temporal ordering, hierarchical composition, and long-range dependencies. We introduce baselines for event trigger detection (ETD), event localization (EL), event argument extraction (EAE), and event relation extraction (ERE). The benchmark is highly challenging for grounded event discovery, with ETD below 8%, EL under 6%, and EAE below 11%. In contrast, ERE is more tractable once events are given, reaching 35.45% F1 zero-shot and 44.42% F1 after fine-tuning.
@inproceedings{asgarov2026nest, title = {NEST: Narrative Event Structures in Time for Long Video Understanding}, author = {Asgarov, Ali and Narasimhan, Kaushik and Sarker, Najibul Haque and Alomari, Hani and Tang, Chia-Wei and Sivakumar, Anushka and Hakim, Zaber Ibn Abdul and Mallampati, Shaurya and Thomas, Chris}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - EMNLP
Investigating Length Bias and Robustness in LVLMs for Multiple-Choice Question AnsweringMd. Atabuzzaman, Hani Alomari, and Chris ThomasIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026Large Vision-Language Models (LVLMs) demonstrate strong performance on multiple-choice question answering (MCQA), yet their reliability under biased conditions remains underexplored. We investigate length bias, the systematic preference for longer answer options, through controlled benchmarks spanning fine-grained visual domains. LVLMs consistently favor longer options even when visually inconsistent, with accuracy dropping 20-40% when correct answers are shorter and increasing 40-70% when longer. This bias persists across model scales. Through layer-wise mechanistic analysis, we reveal that length bias emerges from late-layer suppression of visual evidence in favor of language-side length priors. Blurry-image experiments show that bias intensifies when visual evidence weakens, while abstention tests reveal that LVLMs struggle to refuse incorrect predictions. We propose and evaluate multiple mitigation strategies including Early-Layer Ensemble Decoding (ELED), achieving bias gap reductions of 20-35% through adaptive decoding from pre-override layers. However, residual bias underscores the need for architectural solutions and bias-aware training.
@inproceedings{atabuzzaman2026lengthbias, title = {Investigating Length Bias and Robustness in LVLMs for Multiple-Choice Question Answering}, author = {Atabuzzaman, Md. and Alomari, Hani and Thomas, Chris}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - EMNLP
Reliability Challenges in Diffusion Vision–Language ModelsMd. Atabuzzaman and Chris ThomasIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026Diffusion-based Large Vision-Language Models (dLVLMs) have recently emerged as a compelling alternative to autoregressive (AR) LVLMs, offering advantages in parallel decoding, bidirectional context, and controllable generation. Despite rapid progress, their reliability properties remain largely uncharacterized. We present the first systematic reliability evaluation of hallucination and bias in dLVLMs, benchmarking six diffusion models against competitive AR baselines across four dimensions. Our key findings are: (1) dLVLMs reverse the yes-bias of AR models in binary visual queries; (2) they achieve competitive hallucination rates yet exhibit degraded linguistic quality; (3) they collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias; and (4) they exhibit accuracy collapse in multiple-choice settings when the correct option is shorter than its distractors, associated with a length prior that emerges at the first denoising step. Tokens committed at late denoising steps with low confidence further correlate with hallucinated content, pointing to a mechanistic signal unique to diffusion generation. These patterns vary across model families, suggesting reliability is shaped by the generative paradigm together with training data.
@inproceedings{atabuzzaman2026diffusion, title = {Reliability Challenges in Diffusion Vision–Language Models}, author = {Atabuzzaman, Md. and Thomas, Chris}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - EMNLP
IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective SignalsMd. Atabuzzaman, Christian Alexander, and Chris ThomasIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026Large Vision-Language Models (LVLMs) have achieved strong multimodal performance, yet ensuring the factual correctness of generated content remains challenging. Existing methods that provide statistical guarantees on factuality typically rely on external verifiers or generation-time confidence signals, which introduce auxiliary dependencies or often fail for confident but incorrect outputs. We argue that reliable factuality control can instead be achieved through introspective signals derived from the model itself. We introduce IntroConformal, a training-free Conformal Risk Control (CRC) framework that provides finite-sample, distribution-free factuality guarantees. We first instantiate it with layer-wise semantic stability, a conformity score derived from hidden-state representations, and then propose verification probability, a stronger score capturing the model’s self-administered judgment on claim factuality. Across multiple LVLM architectures, IntroConformal satisfies the conformal risk guarantee while reducing abstention and improving claim-level discrimination over external verifier and decoding-based baselines.
@inproceedings{atabuzzaman2026introconformal, title = {IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals}, author = {Atabuzzaman, Md. and Alexander, Christian and Thomas, Chris}, booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - CVPR
Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization DynamicsNajibul Haque Sarker, Zaber Ibn Abdul Hakim, Ali Asgarov, Chia-Wei Tang, Alvi Md Ishmam, and Chris ThomasIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026Selected as a Highlight paper at CVPR 2026
Fine-tuning has become the default way to adapt powerful foundation models, but this also enables low-cost repurposing for harmful objectives. Existing immunization methods try to optimize local geometry or simulate short attacker horizons, and penalize observed loss drops. However, in practice, downstream tuners run thousands of updates and can often overcome these short-horizon defenses. In this paper, we propose CLAMP (Contractive Long-horizon Attacker Mitigation via Progress-bounding), an immunization method that traps harmful fine-tuning by shaping the attacker’s optimization dynamics rather than only the initial landscape. Our key idea is to make harmful training locally contractive, making each update smaller than the last. This yields a closed-form bound on the attacker’s training beyond the attacker’s simulated training steps. We also introduce a Hessian-free directional curvature penalty, to create adversarial landscapes along harmful descent directions. Our bi-level objective minimizes the attacker’s predicted improvement from train step zero to infinity. Experiments show our method withstands long-horizon fine-tuning across classification, generative, and autoregressive settings, substantially reduces harmful task adaptation, while preserving benign utility and fine-tunability.
@inproceedings{sarker2026immunizing, title = {Immunizing Models Against Harmful Long-Horizon Fine-Tuning via Contractive Optimization Dynamics}, author = {Sarker, Najibul Haque and Hakim, Zaber Ibn Abdul and Asgarov, Ali and Tang, Chia-Wei and Ishmam, Alvi Md and Thomas, Chris}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2026}, publisher = {IEEE}, } - CVPR
Lenses: Toward Polysemous Vision-Language UnderstandingHani Alomari, Ali Asgarov, and Chris ThomasIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026Most vision-language models assume images have a single literal meaning, even though images are polysemous. We propose a retrieval paradigm that models many-to-many relationships between images and text using interpretive lenses and introduce Lenses, a multi-prompt embedding model and dataset for polysemous image-text retrieval. The Lenses dataset contains 105,669 images and 732,405 captions, with each image paired with multiple captions and image-side prompts annotated across five categories: Literal, Figurative, Emotional, Abstract, and Background. Building on a multimodal large language model, the Lenses model uses learned lens tokens to extract lens-specific embeddings for every image and caption and compares these using a lens-masking similarity function with a global fallback that prioritizes same-lens matches while retaining a global pathway. Training uses a category-aware multi-positive contrastive loss and intra-set diversity regularization to align corresponding perspectives while preventing semantic collapse across lenses. We further propose lens-aware evaluation protocols, including category-aware ranking, that better reflect how humans match images and text. Experiments on the Lenses dataset and public benchmarks show that our model outperforms baselines on literal and non-literal retrieval and reduces over-reliance on literal cues.
@inproceedings{alomari2026lenses, title = {Lenses: Toward Polysemous Vision-Language Understanding}, author = {Alomari, Hani and Asgarov, Ali and Thomas, Chris}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2026}, publisher = {IEEE}, } - ECCV
Reasoning-Guided Part-Level Visual Grounding via Reinforcement LearningKazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, Chia-Wei Tang, Anuj Karpatne, and Chris ThomasIn Proceedings of the European Conference on Computer Vision (ECCV), 2026Multimodal large language models (MLLMs) achieve impressive results in zero-shot visual grounding, generating object bounding boxes and segmentation masks from free-form language queries. Despite this success, we show that they remain weak at grounding object parts and fine-grained components. We introduce Object-Part Hierarchical Reflective Grounding (OP-HRG), a prompting paradigm that induces explicit coarse-to-fine decomposition-localizing the parent object before the part-and incorporates model self-critique and conditional refinement. In addition, we propose a novel GRPO reinforcement learning framework with dedicated part-aware rewards. Experiments on PascalPart, PartImageNet, and InstructPart demonstrate improvements in part localization while preserving object-level performance.
@inproceedings{mehrab2026reasoning, title = {Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning}, author = {Mehrab, Kazi Sajeed and Alomari, Hani and Sarker, Najibul Haque and Hakim, Zaber Ibn Abdul and Tang, Chia-Wei and Karpatne, Anuj and Thomas, Chris}, booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)}, year = {2026}, } - ACL
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal ModelsAafiya Hussain, Gaurav Srivastava, Alvi Ishmam, Zaber Hakim, and Chris ThomasIn Proceedings of the Association for Computational Linguistics (ACL), 2026Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal audio-video-language models. We analyze six complementary attack objectives that target different stages of multimodal processing, including audio encoder representations, cross-modal attention, hidden states, and output likelihoods. Across three state-of-the-art models and multiple benchmarks, we show that audio-only perturbations can induce severe multimodal failures, achieving up to 96% attack success rate. We further show that attacks can be successful at low perceptual distortions (LPIPS <= 0.08, SI-SNR >= 0) and benefit more from extended optimization than increased data scale. Transferability across models and encoders remains limited, while speech recognition systems such as Whisper primarily respond to perturbation magnitude, achieving >97% attack success under severe distortion. These results expose a previously overlooked single-modality attack surface in multimodal systems and motivate defenses that enforce cross-modal consistency.
@inproceedings{hussain2026soundbreak, title = {SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models}, author = {Hussain, Aafiya and Srivastava, Gaurav and Ishmam, Alvi and Hakim, Zaber and Thomas, Chris}, booktitle = {Proceedings of the Association for Computational Linguistics (ACL)}, year = {2026}, publisher = {Association for Computational Linguistics}, } - CHI
Designing Multi-Robot Ground Video Sensemaking with Public Safety ProfessionalsPuqi Zhou, Ali Asgarov, Aafiya Hussain, Wonjoon Park, Amit Paudyal, Sameep Shrestha, Chia-Wei Tang, Michael Lighthiser, Michael Hieb, Xuesu Xiao, Chris Thomas, and Sungsoo Ray HongIn Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, 2026Videos from fleets of ground robots can advance public safety by providing scalable situational awareness and reducing professionals’ burden. Yet little is known about how to design and integrate multi-robot videos into public safety workflows. Collaborating with six police agencies, we examined how such videos could be made practical. In Study 1, we present the first testbed for multi-robot ground video sensemaking. The testbed includes 38 events of interest relevant to public safety, a dataset of 20 robot patrol videos (10 day/night pairs) covering EoI types, and 6 design requirements aimed at improving current video sensemaking practices. In Study 2, we built MRVS, a tool that augments multi-robot patrol video streams with a prompt-engineered video understanding model. Participants reported reduced manual workload and greater confidence with LLM-based explanations, while noting concerns about false alarms and privacy. We conclude with implications for designing future multi-robot video sensemaking tools.
@inproceedings{zhou2026designing, title = {Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals}, author = {Zhou, Puqi and Asgarov, Ali and Hussain, Aafiya and Park, Wonjoon and Paudyal, Amit and Shrestha, Sameep and Tang, Chia-Wei and Lighthiser, Michael and Hieb, Michael and Xiao, Xuesu and Thomas, Chris and Hong, Sungsoo Ray}, booktitle = {Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems}, series = {CHI '26}, year = {2026}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, isbn = {9798400722783}, doi = {10.1145/3772318.3790679}, articleno = {1583}, numpages = {22}, } - AAAI
LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained ModelsAlvi Md Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim, and Chris ThomasIn Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-26), 2026Multimodal Large Language Models (MLLMs) like GPT-4V, Gemini, LLaVA-NeXT, and Idefics have made significant advancements in visual-language understanding and generation, particularly for single-image tasks such as VQA. A few open-source models, such as Mantis and VILA, extend these capabilities to multi-image inputs, enabling coreference, comparison, reasoning, and temporal understanding. Despite their remarkable performance, the adversarial robustness of multi-image MLLMs remains unexplored. In practice, attackers typically only have access to public pretrained models and lack knowledge of downstream models, and gradient-based white-box methods generate instance-specific perturbations that generalize poorly and require costly re-optimization for new inputs. This paper introduces LAMP, a black-box method for learning Universal Adversarial Perturbations (UAPs) targeting multi-image MLLMs. LAMP applies an attention-based constraint that prevents the model from effectively aggregating information across images. LAMP also introduces a novel cross-image contagious constraint that forces perturbed tokens to influence clean tokens, spreading adversarial effects without requiring all inputs to be modified. Additionally, an index-attention suppression loss enables a robust position-invariant attack. Experimental results show that LAMP outperforms SOTA baselines and achieves the highest attack success rates across multiple vision-language tasks and models.
@inproceedings{ishmam2026lamp, title = {LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models}, author = {Ishmam, Alvi Md and Sarker, Najibul Haque and Hakim, Zaber Ibn Abdul and Thomas, Chris}, booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence (AAAI-26)}, year = {2026}, publisher = {AAAI Press}, }
2025
- EMNLP
Flexible-length Text Infilling for Discrete Diffusion ModelsAndrew Zhang, Anushka Sivakumar, Chia-Wei Tang, and Chris ThomasIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025Discrete diffusion models are a new class of text generation models that offer advantages such as bidirectional context, parallelizable generation, and flexible prompting compared to autoregressive models. However, a critical limitation has been the inability to perform flexible-length or flexible-position text infilling without access to ground-truth positional data. We introduce DDOT (Discrete Diffusion with Optimal Transport Position Coupling), a discrete diffusion model that overcomes this limitation by jointly denoising token values and token positions using a novel sample-level optimal transport coupling. This coupling preserves relative token ordering while dynamically adjusting the positions and lengths of infilled segments. DDOT is orthogonal to existing discrete text diffusion methods and is compatible with various pretrained text denoisers. On text-infilling benchmarks such as One-Billion-Word and Yelp, DDOT outperforms naive diffusion baselines and achieves performance on par with state-of-the-art non-autoregressive models, while improving training efficiency and prompting flexibility.
@inproceedings{zhang2025flexible, title = {Flexible-length Text Infilling for Discrete Diffusion Models}, author = {Zhang, Andrew and Sivakumar, Anushka and Tang, Chia-Wei and Thomas, Chris}, booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing}, year = {2025}, address = {Suzhou, China}, publisher = {Association for Computational Linguistics}, } - EMNLP
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language ModelsMd. Atabuzzaman, Ali Asgarov, and Chris ThomasIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025Large Vision-Language Models (LVLMs) have achieved strong performance on vision-language tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored. In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options. We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model’s output. Our method mitigates bias without retraining and is compatible with frozen LVLMs. Extensive experiments across several state-of-the-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings. This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning. Datasets and code are available at: https://github.com/Atabuzzaman/Selection-Bias-of-LVLMs
@inproceedings{atabuzzaman2025mcqa, title = {Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models}, author = {Atabuzzaman, Md. and Asgarov, Ali and Thomas, Chris}, booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing}, year = {2025}, address = {Suzhou, China}, publisher = {Association for Computational Linguistics}, } - EMNLPZero-Shot Fine-Grained Image Classification Using Large Vision-Language ModelsMd. Atabuzzaman, Andrew Zhang, and Chris ThomasIn Findings of the Association for Computational Linguistics: EMNLP 2025, 2025
Large Vision-Language Models (LVLMs) have demonstrated impressive performance on vision-language reasoning tasks. However, their potential for zero-shot fine-grained image classification, a challenging task requiring precise differentiation between visually similar categories, remains underexplored. We present a novel method that transforms zero-shot fine-grained image classification into a visual question-answering framework, leveraging LVLMs’ comprehensive understanding capabilities rather than relying on direct class name generation. We enhance model performance through a novel attention intervention technique. We also address a key limitation in existing datasets by developing more comprehensive and precise class description benchmarks. We validate the effectiveness of our method through extensive experimentation across multiple fine-grained image classification benchmarks. Our proposed method consistently outperforms the current state-of-the-art (SOTA) approach, demonstrating both the effectiveness of our method and the broader potential of LVLMs for zero-shot fine-grained classification tasks.
@inproceedings{atabuzzaman2025zeroshot, title = {Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models}, author = {Atabuzzaman, Md. and Zhang, Andrew and Thomas, Chris}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025}, year = {2025}, address = {Suzhou, China}, publisher = {Association for Computational Linguistics}, } - EMNLPSteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language ModelsAnushka Sivakumar, Andrew Zhang, Zaber Ibn Abdul Hakim, and Chris ThomasIn Findings of the Association for Computational Linguistics: EMNLP 2025, 2025
This work introduces SteerVLM, a lightweight steering module designed to guide Vision Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the language modality with image context. This provides fine-grained, inference-time control over complex output semantics without modifying model weights while preserving performance on off-target tasks. Our steering module requires learning parameters equal to 0.14% of the original VLM’s size. Additionally, our steering module gains model control via dimension-wise activation modulation and adaptive layer-wise steering without requiring pre-extracted static vectors or manual tuning of intervention points. Furthermore, we introduce VNIA (Visual Narrative Intent Alignment), a multimodal dataset specifically created to facilitate the development and evaluation of VLM steering techniques. Our method outperforms existing intervention techniques on steering and hallucination mitigation benchmarks for VLMs and proposes a robust solution for multimodal model control through activation engineering.
@inproceedings{sivakumar2025steervlm, title = {SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models}, author = {Sivakumar, Anushka and Zhang, Andrew and Hakim, Zaber Ibn Abdul and Thomas, Chris}, booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025}, year = {2025}, address = {Suzhou, China}, publisher = {Association for Computational Linguistics}, } - ACL
Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal RetrievalHani Alomari, Anushka Sivakumar, Andrew Zhang, and Chris ThomasIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025Cross-modal image-text retrieval is challenging because of the diverse possible associations between content from different modalities. Traditional methods learn a single-vector embedding to represent semantics of each sample, but struggle to capture nuanced and diverse relationships that can exist across modalities. Set-based approaches, which represent each sample with multiple embeddings, offer a promising alternative, as they can capture richer and more diverse relationships. In this paper, we show that, despite their promise, these set-based representations continue to face issues including sparse supervision and set collapse, which limits their effectiveness. To address these challenges, we propose Maximal Pair Assignment Similarity to optimize one-to-one matching between embedding sets which preserve semantic diversity within the set. We also introduce two loss functions to further enhance the representations: Global Discriminative Loss to enhance distinction among embeddings, and Intra-Set Divergence Loss to prevent collapse within each set. Our method achieves state-of-the-art performance on MS-COCO and Flickr30k without relying on external data.
@inproceedings{alomari2025maximal, title = {Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval}, author = {Alomari, Hani and Sivakumar, Anushka and Zhang, Andrew and Thomas, Chris}, booktitle = {Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL)}, year = {2025}, address = {Vienna, Austria}, pages = {31769--31785}, publisher = {Association for Computational Linguistics}, } - CVPRW
Real-Time Ultra-Fine-Grained Surgical Instrument ClassificationMd. Atabuzzaman, Gino DiMatteo, Hani Alomari, Chia-Wei Tang, Connor Hale, Adam E. Goode, David Ryan King, and Chris ThomasIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Fine-Grained Visual Categorization (FGVC) Workshop, 2025Accurate classification of ultra-fine-grained surgical instruments can significantly reduce the rate of canceled or postponed surgical procedures and improve a hospital’s overall operational efficiency. However, accurately classifying these instruments is challenging due to the vast number of surgical instruments in a hospital’s Central Sterile Services Department (CSSD) and their ultra-fine-grained distinctions. To address this challenge and assist CSSD technicians, we propose a real-time ultra-fine-grained surgical instrument classification system. Our system consists of a unique open-environment image acquisition platform and multi-view CNN and transformer-based architectures to capture and classify multi-view images of instruments in real-time. We train models on images from three globally recognized surgical trays: Eye Vitrectomy, Major Laparotomy, and Minor Laparotomy, encompassing 95 distinct classes. We evaluate our system in real-time and on image-based datasets, demonstrating state-of-the-art (SoTA) performance. A user study conducted after deployment in a hospital CSSD reveals that the system significantly improves workflow efficiency, streamlining CSSD operations.
@inproceedings{atabuzzaman2025realtime, title = {Real-Time Ultra-Fine-Grained Surgical Instrument Classification}, author = {Atabuzzaman, Md. and DiMatteo, Gino and Alomari, Hani and Tang, Chia-Wei and Hale, Connor and Goode, Adam E. and King, David Ryan and Thomas, Chris}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Fine-Grained Visual Categorization (FGVC) Workshop}, year = {2025}, pages = {2070--2079}, publisher = {IEEE}, } - NeurIPSW
Model Immunization by Trapping Harmful FinetuningNajibul Haque Sarker, Zaber Ibn Abdul Hakim, Alvi Md Ishmam, Chia-Wei Tang, and Chris ThomasIn NeurIPS 2025 Workshop on Lock-LLM: Prevent Unauthorized Knowledge Use from LLMs, 2025Model immunization is a new technique of protecting models against downstream harmful fine-tuning while remaining useful on intended tasks. Prior works utilize condition number based regularizers to ill-condition the optimization landscape for harmful tasks. However, the induced protection does not guarantee that immunization will persist. In this work, we introduce the novel concept of creating a trap in the landscape, so that harmful finetuning optimization will be trapped in an unoptimized minima. We propose a geometry-aware trap-inducing objective, which limits multi-step harmful loss reduction to the expected local geometry-based loss. Furthermore, to properly evaluate immunization retainment, we introduce an extrinsic metric, Relative Fine-Tuning Deviation (RFD). Across multiple pretrained backbones and datasets, we show our method increases resistance to harmful adaptation and preserves primary-task accuracy, outperforming curvature-only baselines on RFD while remaining competitive on standard utility metrics.
@inproceedings{sarker2025model, title = {Model Immunization by Trapping Harmful Finetuning}, author = {Sarker, Najibul Haque and Hakim, Zaber Ibn Abdul and Ishmam, Alvi Md and Tang, Chia-Wei and Thomas, Chris}, booktitle = {NeurIPS 2025 Workshop on Lock-LLM: Prevent Unauthorized Knowledge Use from LLMs}, year = {2025}, } - arXivPAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language ModelKazi Hasan Ibn Arif, Sajib Acharjee Dip, Khizar Hussain, Lang Zhang, and Chris ThomasarXiv preprint arXiv:2501.12206, 2025
Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities in understanding and describing visual content, achieving state-of-the-art performance across various vision-language tasks. However, these models often generate descriptions containing objects or details that are absent in the input image, a phenomenon commonly known as hallucination. Our work investigates the key reasons behind this issue by analyzing the pattern of self-attention in transformer layers. We find that hallucinations often arise from the progressive weakening of attention weight to visual tokens in the deeper layers of the LLM. Some previous works naively boost the attention of all visual tokens to mitigate this issue, resulting in suboptimal hallucination reduction. To address this, we identify two critical sets of visual tokens that facilitate the transfer of visual information from the vision encoder to the LLM. Local tokens encode grounded information about objects present in an image, while summary tokens capture the overall aggregated representation of the image. Importantly, these two sets of tokens require different levels of weight enhancement. To this end, we propose PAINT (Paying Attention to INformed Tokens), a plug-and-play framework that intervenes in the self-attention mechanism of the LLM, selectively boosting the attention weights of local and summary tokens with experimentally learned margins. Evaluation on the MSCOCO image captioning dataset demonstrate that our approach reduces hallucination rates by up to 62.3% compared to baseline models while maintaining accuracy. Code is available at https://github.com/hasanar1f/PAINT
@article{arif2025paint, title = {PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model}, author = {Arif, Kazi Hasan Ibn and Dip, Sajib Acharjee and Hussain, Khizar and Zhang, Lang and Thomas, Chris}, journal = {arXiv preprint arXiv:2501.12206}, year = {2025}, } - WACVAdvancing chart question answering with robust chart component recognitionHanwen Zheng, Sijia Wang, Chris Thomas, and Lifu HuangIn 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025
Chart comprehension presents significant challenges for machine learning models due to the diverse and intricate shapes of charts. Existing multimodal methods often overlook these visual features or fail to integrate them effectively for chart question answering (ChartQA). To address this, we introduce CHARTFORMER, a unified framework that enhances chart component recognition by accurately identifying and classifying components such as bars, lines, pies, titles, legends, and axes. Additionally, we propose a novel Question-guided Deformable Co-Attention (QDCAt) mechanism, which fuses chart features encoded by CHARTFORMER with the given question, leveraging the question’s guidance to ground the correct answer. Extensive experiments demonstrate that the proposed approaches significantly outperform baseline models in chart component recognition and ChartQA tasks, achieving improvements of 3.2% in mAP and 15.4% in accuracy, respectively. These results underscore the robustness of our solution for detailed visual data interpretation across various applications.
@inproceedings{zheng2025advancing, title = {Advancing chart question answering with robust chart component recognition}, author = {Zheng, Hanwen and Wang, Sijia and Thomas, Chris and Huang, Lifu}, booktitle = {2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)}, pages = {5741--5750}, year = {2025}, organization = {IEEE}, }
2024
- NeurIPSWENTER: Event Based Interpretable Reasoning for VideoQAHammad Ayyubi, Junzhang Liu, Ali Asgarov, Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Zhecan Wang, Chia-Wei Tang, Hani Alomari, Md Atabuzzaman, Xudong Lin, Naveen Reddy Dyava, Shih-Fu Chang, and Chris ThomasIn Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2024, 2024
Selected as a Spotlight at the NeurIPS 2024 Multimodal Algorithmic Reasoning (MAR) Workshop
In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where video events form the nodes and event-event relationships (temporal/causal/hierarchical) form the edges. This structured representation offers many benefits: 1) Interpretable VideoQA via generated code that parses event-graph; 2) Incorporation of contextual visual information in the reasoning process (code generation) via event graphs; 3) Robust VideoQA via Hierarchical Iterative Update of the event graphs. Existing interpretable VideoQA systems are often top-down, disregarding low-level visual information in the reasoning plan generation, and are brittle. While bottom-up approaches produce responses from visual data, they lack interpretability. Experimental results on NExT-QA, IntentQA, NExT-GQA, and STAR demonstrate that not only does our method outperform existing top-down approaches while obtaining competitive performance against bottom-up approaches, but more importantly, it offers superior interpretability and explainability in the reasoning process.
@inproceedings{ayyubi2025enter, title = {ENTER: Event Based Interpretable Reasoning for VideoQA}, author = {Ayyubi, Hammad and Liu, Junzhang and Asgarov, Ali and Hakim, Zaber Ibn Abdul and Sarker, Najibul Haque and Wang, Zhecan and Tang, Chia-Wei and Alomari, Hani and Atabuzzaman, Md and Lin, Xudong and Dyava, Naveen Reddy and Chang, Shih-Fu and Thomas, Chris}, booktitle = {Multimodal Algorithmic Reasoning (MAR) Workshop, NeurIPS 2024}, year = {2024}, } - EMNLP
M3D: MultiModal MultiDocument Fine-Grained Inconsistency DetectionChia-Wei Tang, Ting-Chih Chen, Kiet Nguyen, Kazi Sajeed Mehrab, Alvi Ishmam, and Chris ThomasIn Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024Fact-checking claims is a highly laborious task that involves understanding how each factual assertion within the claim relates to a set of trusted source materials. Existing approaches make sample-level predictions but fail to identify the specific aspects of the claim that are troublesome and the specific evidence relied upon. In this paper, we introduce a method and new benchmark for this challenging task. Our method predicts the fine-grained logical relationship of each aspect of the claim from a set of multimodal documents, which include text, image(s), video(s), and audio(s). We also introduce a new benchmark (M3DC) of claims requiring multimodal multidocument reasoning, which we construct using a novel claim synthesis technique. Experiments show that our approach outperforms other models on this challenging task on two benchmarks while providing finer-grained predictions, explanations, and evidence.
@inproceedings{tang2024m3d, title = {M3D: MultiModal MultiDocument Fine-Grained Inconsistency Detection}, author = {Tang, Chia-Wei and Chen, Ting-Chih and Nguyen, Kiet and Mehrab, Kazi Sajeed and Ishmam, Alvi and Thomas, Chris}, booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing}, pages = {22270--22293}, year = {2024}, } - ACMDetecting Multimodal Situations with Insufficient Context and Abstaining from Baseless PredictionsJunzhang Liu, Zhecan Wang, Hammad Ayyubi, Haoxuan You, Chris Thomas, Rui Sun, Shih-Fu Chang, and Kai-Wei ChangIn Proceedings of the 32nd ACM International Conference on Multimedia, 2024
Despite the widespread adoption of Vision-Language Understanding (VLU) benchmarks such as VQA v2, OKVQA, A-OKVQA, GQA, VCR, SWAG, and VisualCOMET, our analysis reveals a pervasive issue affecting their integrity: these benchmarks contain samples where answers rely on assumptions unsupported by the provided context. Training models on such data foster biased learning and hallucinations as models tend to make similar unwarranted assumptions. To address this issue, we collect contextual data for each sample whenever available and train a context selection module to facilitate evidence-based model predictions. Strong improvements across multiple benchmarks demonstrate the effectiveness of our approach. Further, we develop a general-purpose Context-AwaRe Abstention (CARA) detector to identify samples lacking sufficient context and enhance model accuracy by abstaining from responding if the required context is absent. CARA exhibits generalization to new benchmarks it wasn’t trained on, underscoring its utility for future VLU benchmarks in detecting or cleaning samples with inadequate context. Finally, we curate a Context Ambiguity and Sufficiency Evaluation (CASE) set to benchmark the performance of insufficient context detectors. Overall, our work represents a significant advancement in ensuring that vision-language models generate trustworthy and evidence-based outputs in complex real-world scenarios.
@inproceedings{liu2024detecting, title = {Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions}, author = {Liu, Junzhang and Wang, Zhecan and Ayyubi, Hammad and You, Haoxuan and Thomas, Chris and Sun, Rui and Chang, Shih-Fu and Chang, Kai-Wei}, booktitle = {Proceedings of the 32nd ACM International Conference on Multimedia}, pages = {8402--8411}, year = {2024}, } - NeurIPS
Journeybench: A challenging one-stop vision-language understanding benchmark of generated imagesZhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun, Wenhao Li, Md. Atabuzzaman, Hammad Ayyubi, Haoxuan You, Alvi Md Ishmam, Kai-Wei Chang, Shih-Fu Chang, and Chris ThomasIn Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on these benchmarks does not necessarily correlate with strong visual understanding. In this paper, we release JourneyBench, a comprehensive human-annotated benchmark of generated images designed to assess the model’s fine-grained multimodal reasoning abilities across five tasks: complementary multimodal chain of thought, multi-image VQA, imaginary image captioning, VQA with hallucination triggers, and fine-grained retrieval with sample-specific distractors. Unlike existing benchmarks, JourneyBench explicitly requires fine-grained multimodal reasoning in unusual imaginary scenarios where language bias and holistic image gist are insufficient. We benchmark state-of-the-art models on JourneyBench and analyze performance along a number of fine-grained dimensions. Results across all five tasks show that JourneyBench is exceptionally challenging for even the best models, indicating that models’ visual reasoning abilities are not as strong as they first appear. We discuss the implications of our findings and propose avenues for further research.
@inproceedings{wang2024journeybench, title = {Journeybench: A challenging one-stop vision-language understanding benchmark of generated images}, author = {Wang, Zhecan and Liu, Junzhang and Tang, Chia-Wei and Alomari, Hani and Sivakumar, Anushka and Sun, Rui and Li, Wenhao and Atabuzzaman, Md. and Ayyubi, Hammad and You, Haoxuan and Ishmam, Alvi Md and Chang, Kai-Wei and Chang, Shih-Fu and Thomas, Chris}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track}, year = {2024}, } - ACL
MetaSumPerceiver: Multimodal Multi-Document Evidence Summarization for Fact-CheckingTing-Chih Chen, Chia-Wei Tang, and Christopher ThomasIn Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024Fact-checking real-world claims often requires reviewing multiple multimodal documents in order to assess the claim’s truthfulness, a highly laborious and time-consuming task. In this paper, we present a summarization model crafted to generate claim-specific summaries useful for fact-checking from multimodal multi-document datasets. The model takes inputs in the form of documents, images, and a claim, with the objective of assisting in fact-checking tasks. We introduce a dynamic perceiver-based model that is able to handle inputs from multiple modalities of arbitrary lengths. To train our model, we leverage a novel reinforcement learning-based entailment objective in order to generate summaries that provide evidence distinguishing between different truthfulness labels. To assess the efficacy of our approach, we conduct experiments on both an existing benchmark as well as a new dataset of multi-document claims which we contribute. Our approach outperforms the SOTA approach by 4.6% in the claim verification task on the MOCHEG dataset and demonstrates strong performance on our new Multi-News-Fact-Checking dataset.
@inproceedings{chen-etal-2024-metasumperceiver, title = {{M}eta{S}um{P}erceiver: Multimodal Multi-Document Evidence Summarization for Fact-Checking}, author = {Chen, Ting-Chih and Tang, Chia-Wei and Thomas, Christopher}, editor = {Ku, Lun-Wei and Martins, Andre and Srikumar, Vivek}, booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, year = {2024}, address = {Bangkok, Thailand}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2024.acl-long.474}, pages = {8742--8757}, } - CVPR
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge AlignmentAlvi Md Ishmam and Christopher ThomasIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024In recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web for training also makes these models vulnerable to potential security threats, such as backdooring and poisoning attacks. In this paper, we propose a method for mitigating such attacks on contrastively trained vision-language models. Our approach leverages external knowledge extracted from a language model to prevent models from learning correlations between image regions which lack strong alignment with external knowledge. We do this by imposing constraints to enforce that attention paid by the model to visual regions is proportional to the alignment of those regions with external knowledge. We conduct extensive experiments using a variety of recent backdooring and poisoning attacks on multiple datasets and architectures. Our results clearly demonstrate that our proposed approach is highly effective at defending against such attacks across multiple settings, while maintaining model utility and without requiring any changes at inference time.
@inproceedings{ishmam2024semantic, title = {Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment}, author = {Ishmam, Alvi Md and Thomas, Christopher}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, pages = {24820--24830}, year = {2024}, } -
Beyond Grounding: Extracting Fine-Grained Event Hierarchies across ModalitiesHammad Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh, Long Chen, Yulei Niu, Xudong Lin, Xuande Feng, Jaywon Koo, Sounak Ray, and othersIn Proceedings of the AAAI Conference on Artificial Intelligence, 2024Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if events across textual and visual (video) domains are identical (via grounding) and thus, on the same semantic level. However, grounding fails to capture the intricate cross-event relations that exist due to the same events being referred to on many semantic levels. For example, the abstract event of "war" manifests at a lower semantic level through subevents "tanks firing" (in video) and airplane "shot" (in text), leading to a hierarchical, multimodal relationship between the events. In this paper, we propose the task of extracting event hierarchies from multimodal (video and text) data to capture how the same event manifests itself in different modalities at different semantic levels. This reveals the structure of events and is critical to understanding them. To support research on this task, we introduce the Multimodal Hierarchical Events (MultiHiEve) dataset. Unlike prior video-language datasets, MultiHiEve is composed of news video-article pairs, which makes it rich in event hierarchies. We densely annotate a part of the dataset to construct the test benchmark. We show the limitations of state-of-the-art unimodal and multimodal baselines on this task. Further, we address these limitations via a new weakly supervised model, leveraging only unannotated video-article pairs from MultiHiEve. We perform a thorough evaluation of our proposed method which demonstrates improved performance on this task and highlight opportunities for future research.
@inproceedings{ayyubi2024beyond, title = {Beyond Grounding: Extracting Fine-Grained Event Hierarchies across Modalities}, author = {Ayyubi, Hammad and Thomas, Christopher and Chum, Lovish and Lokesh, Rahul and Chen, Long and Niu, Yulei and Lin, Xudong and Feng, Xuande and Koo, Jaywon and Ray, Sounak and others}, booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence}, volume = {38}, number = {16}, pages = {17664--17672}, year = {2024}, }
2023
- ACLEnhanced Chart Understanding via Visual Language Pre-training on Plot Table PairsMingyang Zhou, Yi Fung, Long Chen, Christopher Thomas, Heng Ji, and Shih-Fu ChangIn Findings of the Association for Computational Linguistics: ACL 2023, 2023
Building cross-model intelligence that can understand charts and communicate the salient information hidden behind them is an appealing challenge in the vision and language (V+L) community. The capability to uncover the underlined table data of chart figures is a critical key to automatic chart understanding. We introduce ChartT5, a V+L model that learns how to interpret table information from chart images via cross-modal pre-training on plot table pairs. Specifically, we propose two novel pre-training objectives: Masked Header Prediction (MHP) and Masked Value Prediction (MVP) to facilitate the model with different skills to interpret the table information. We have conducted extensive experiments on chart question answering and chart summarization to verify the effectiveness of the proposed pre-training strategies. In particular, on the ChartQA benchmark, our ChartT5 outperforms the state-of-the-art non-pretraining methods by over 8% performance gains.
@inproceedings{zhou-etal-2023-enhanced, title = {Enhanced Chart Understanding via Visual Language Pre-training on Plot Table Pairs}, author = {Zhou, Mingyang and Fung, Yi and Chen, Long and Thomas, Christopher and Ji, Heng and Chang, Shih-Fu}, editor = {Rogers, Anna and Boyd-Graber, Jordan and Okazaki, Naoaki}, booktitle = {Findings of the Association for Computational Linguistics: ACL 2023}, year = {2023}, address = {Toronto, Canada}, publisher = {Association for Computational Linguistics}, url = {https://aclanthology.org/2023.findings-acl.85}, doi = {10.18653/v1/2023.findings-acl.85}, pages = {1314--1326}, }
2022
- ECCV
Fine-Grained Visual EntailmentChristopher Thomas, Yipeng Zhang, and Shih-Fu ChangIn Proceedings of the European Conference on Computer Vision, 2022@inproceedings{thomas2022fine, title = {Fine-Grained Visual Entailment}, author = {Thomas, Christopher and Zhang, Yipeng and Chang, Shih-Fu}, booktitle = {Proceedings of the European Conference on Computer Vision}, pages = {398--416}, year = {2022}, } - Community implications for gun violence prevention during co-occurring pandemics; a qualitative and computational analysis studyDesmond U. Patton, Nathan Aguilar, Aviv Y. Landau, Chris Thomas, Rachel Kagan, Tianai Ren, Eric Stoneberg, Timothy Wang, Daniel Halmos, Anish Saha, Amith Ananthram, and Kathleen McKeownPreventive Medicine, 2022
This study provides insight into New York City residents’ perceptions about violence after the outbreak of Coronavirus disease (COVID-19) based on information from communities in New York City Housing Authority (NYCHA) buildings. In this novel analysis, we used focus group and social media data to confirm or reject findings from qualitative interviews. We first used data from 69 in-depth, semi-structured interviews with low-income residents and community stakeholders to further explore how violence impacts New York City’s low-income residents of color, as well as the role of city government in providing tangible support for violence prevention during co-occurring health (COVID-19) and social (anti-Black racism) pandemics. Residents described how COVID-19 and the Black Lives Matter movement impacted safety in their communities while offering direct recommendations to improve safety. Residents also shared recommendations that indirectly improve community safety by addressing long term systemic issues. As the recruitment of interviewees was concluding, researchers facilitated two focus groups with 38 interviewees to discuss similar topics. In order to assess the degree to which the themes discovered in our qualitative interviews were shared by the broader community, we developed an integrative community data science study which leveraged natural language processing and computer vision techniques to study text and images on public social media data of 12 million tweets generated by residents. We joined computational methods with qualitative analysis through a social work lens and design justice principles to most accurately and holistically analyze the community perceptions of gun violence issues and potential prevention strategies. Findings indicate valuable community-based insights that elucidate how the co-occurring pandemics impact residents’ experiences of gun violence and provide important implications for gun violence prevention in a digital era.
@article{PATTON2022107263, title = {Community implications for gun violence prevention during co-occurring pandemics; a qualitative and computational analysis study}, journal = {Preventive Medicine}, volume = {165}, pages = {107263}, year = {2022}, issn = {0091-7435}, doi = {10.1016/j.ypmed.2022.107263}, url = {https://www.sciencedirect.com/science/article/pii/S0091743522003127}, author = {Patton, Desmond U. and Aguilar, Nathan and Landau, Aviv Y. and Thomas, Chris and Kagan, Rachel and Ren, Tianai and Stoneberg, Eric and Wang, Timothy and Halmos, Daniel and Saha, Anish and Ananthram, Amith and McKeown, Kathleen}, keywords = {Gun violence, COVID-19, Black lives matter, Defund the police, Social media, Qualitative and computational analysis}, } - CVPRW
Emphasizing Complementary Samples for Non-Literal Cross-Modal RetrievalChristopher Thomas and Adriana KovashkaIn Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2022@inproceedings{thomas2022emphasizing, title = {Emphasizing Complementary Samples for Non-Literal Cross-Modal Retrieval}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops}, pages = {4632--4641}, year = {2022}, } - TPAMILearning to Overcome Noise in Weak Caption Supervision for Object DetectionMesut Erhan Unal, Keren Ye, Mingda Zhang, Christopher Thomas, Adriana Kovashka, Wei Li, Danfeng Qin, and Jesse BerentIEEE Transactions on Pattern Analysis and Machine Intelligence, 2022
@article{unal2022learning, title = {Learning to Overcome Noise in Weak Caption Supervision for Object Detection}, author = {Unal, Mesut Erhan and Ye, Keren and Zhang, Mingda and Thomas, Christopher and Kovashka, Adriana and Li, Wei and Qin, Danfeng and Berent, Jesse}, journal = {IEEE Transactions on Pattern Analysis and Machine Intelligence}, year = {2022}, publisher = {IEEE} } - EMNLPWeakly-Supervised Temporal Article GroundingLong Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, and Shih-Fu ChangIn Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP2022), 2022
@inproceedings{chen2022weakly, title = {Weakly-Supervised Temporal Article Grounding}, author = {Chen, Long and Niu, Yulei and Chen, Brian and Lin, Xudong and Han, Guangxing and Thomas, Christopher and Ayyubi, Hammad and Ji, Heng and Chang, Shih-Fu}, booktitle = {Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP2022)}, year = {2022} }
2021
- ACL
InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News DetectionYi Fung, Christopher Thomas, Revanth Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, and Avi SilIn Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021Selected for an Oral presentation at ACL 2021
@inproceedings{fung2021infosurgeon, title = {InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection}, author = {Fung, Yi and Thomas, Christopher and Reddy, Revanth and Polisetty, Sandeep and Ji, Heng and Chang, Shih-Fu and McKeown, Kathleen and Bansal, Mohit and Sil, Avi}, booktitle = {Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL)}, pages = {1683--1698}, year = {2021}, } - IJCVPredicting Visual Political Bias Using Webly Supervised Data and an Auxiliary TaskChristopher Thomas and Adriana KovashkaInternational Journal of Computer Vision, 2021
@article{thomas2021predicting, doi = {10.1007/s11263-021-01506-3}, title = {Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task}, author = {Thomas, Christopher and Kovashka, Adriana}, journal = {International Journal of Computer Vision}, volume = {129}, number = {11}, pages = {2978--3003}, year = {2021}, publisher = {Springer US} } - EMNLPJoint Multimedia Event Extraction from Video and ArticleBrian Chen, Xudong Lin, Christopher Thomas, Manling Li, Shoya Yoshida, Lovish Chum, Heng Ji, and Shih-Fu ChangIn Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP2021) Findings, 2021
@inproceedings{chen2021joint, title = {Joint Multimedia Event Extraction from Video and Article}, author = {Chen, Brian and Lin, Xudong and Thomas, Christopher and Li, Manling and Yoshida, Shoya and Chum, Lovish and Ji, Heng and Chang, Shih-Fu}, booktitle = {Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP2021) Findings}, pages = {74--88}, year = {2021} }
2020
- ECCV
Preserving Semantic Neighborhoods for Robust Cross-modal RetrievalChristopher Thomas and Adriana KovashkaIn Proceedings of the European Conference on Computer Vision (ECCV), 2020@inproceedings{thomas2020preserving, title = {Preserving Semantic Neighborhoods for Robust Cross-modal Retrieval}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)}, year = {2020}, } - arXivLearning to Transfer Visual Effects from Videos to ImagesChristopher Thomas, Yale Song, and Adriana KovashkaarXiv preprint arXiv:2012.01642, 2020
@article{thomas2020learning, title = {Learning to Transfer Visual Effects from Videos to Images}, author = {Thomas, Christopher and Song, Yale and Kovashka, Adriana}, journal = {arXiv preprint arXiv:2012.01642}, year = {2020} }
2019
- NeurIPS
Predicting the politics of an image using webly supervised dataChristopher Thomas and Adriana KovashkaIn Advances in Neural Information Processing Systems (NeurIPS 2019), 2019@inproceedings{thomas2019predicting, title = {Predicting the politics of an image using webly supervised data}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS 2019)}, year = {2019}, }
2018
- BMVC
Persuasive faces: generating faces in advertisementsChristopher Thomas and Adriana KovashkaIn Proceedings of the British Machine Vision Conference, 2018@inproceedings{thomas2018persuasive, title = {Persuasive faces: generating faces in advertisements}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Proceedings of the British Machine Vision Conference}, year = {2018}, } - ACCVArtistic object recognition by unsupervised style adaptationChristopher Thomas and Adriana KovashkaIn Asian Conference on Computer Vision, 2018
@inproceedings{thomas2018artistic, title = {Artistic object recognition by unsupervised style adaptation}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Asian Conference on Computer Vision}, pages = {460--476}, year = {2018}, organization = {Springer, Cham} }
2017
- CVPR
Automatic understanding of image and video advertisementsZaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, and Adriana KovashkaIn Proceedings of the IEEE conference on computer vision and pattern recognition, 2017Selected as a Spotlight at CVPR 2017
@inproceedings{hussain2017automatic, title = {Automatic understanding of image and video advertisements}, author = {Hussain, Zaeem and Zhang, Mingda and Zhang, Xiaozhong and Ye, Keren and Thomas, Christopher and Agha, Zuha and Ong, Nathan and Kovashka, Adriana}, booktitle = {Proceedings of the IEEE conference on computer vision and pattern recognition}, pages = {1705--1715}, year = {2017}, }
2016
- CVPR
Seeing Behind the Camera: Identifying the Authorship of a PhotographChristopher Thomas and Adriana KovashkaIn Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016@inproceedings{thomas2016seeing, title = {Seeing Behind the Camera: Identifying the Authorship of a Photograph}, author = {Thomas, Christopher and Kovashka, Adriana}, booktitle = {Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition}, pages = {3494--3501}, year = {2016}, } - arXivOpensalicon: An open source implementation of the salicon saliency modelChristopher ThomasarXiv preprint arXiv:1606.00110, 2016
@article{thomas2016opensalicon, title = {Opensalicon: An open source implementation of the salicon saliency model}, author = {Thomas, Christopher}, journal = {arXiv preprint arXiv:1606.00110}, year = {2016} } - CVPRWA Visual Attention Algorithm Designed for Coupled Oscillator AccelerationChristopher Thomas, Adriana Kovashka, Donald Chiarulli, and Steven LevitanIn Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016
@inproceedings{thomas2016visual, title = {A Visual Attention Algorithm Designed for Coupled Oscillator Acceleration}, author = {Thomas, Christopher and Kovashka, Adriana and Chiarulli, Donald and Levitan, Steven}, booktitle = {Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops}, pages = {10--18}, year = {2016} }
2015
- arXivHand Posture’s Effect on Touch Screen Text Input Behaviors: A Touch Area Based StudyChristopher Thomas and Brandon JenningsarXiv preprint arXiv:1504.02134, 2015
@article{thomas2015hand, title = {Hand Posture's Effect on Touch Screen Text Input Behaviors: A Touch Area Based Study}, author = {Thomas, Christopher and Jennings, Brandon}, journal = {arXiv preprint arXiv:1504.02134}, year = {2015} } - SEKEApplication of Slow Intelligence Framework for Smart Pet Care System DesignShi-Kuo Chang, Wen-Hui Chen, Wen-Chyi Lin, and Christopher Lee ThomasIn Software Engineering and Knowledge Engineering (SEKE 2015), 2015
@inproceedings{chang2015application, title = {Application of Slow Intelligence Framework for Smart Pet Care System Design}, author = {Chang, Shi-Kuo and Chen, Wen-Hui and Lin, Wen-Chyi and Thomas, Christopher Lee}, booktitle = {Software Engineering and Knowledge Engineering (SEKE 2015)}, pages = {74--79}, year = {2015} }
2014
- INLGTBI-Doc: Generating patient & clinician reports from brain imaging dataPamela Jordan, Nancy Green, Christopher Thomas, and Susan HolmIn Proceedings of the 8th International Natural Language Generation Conference (INLG), 2014
@inproceedings{jordan2014tbi, title = {TBI-Doc: Generating patient \& clinician reports from brain imaging data}, author = {Jordan, Pamela and Green, Nancy and Thomas, Christopher and Holm, Susan}, booktitle = {Proceedings of the 8th International Natural Language Generation Conference (INLG)}, pages = {143--146}, year = {2014} } - Student Response AnalysisSean Myers, Timothy Parenti, and Chris Thomas2014
@misc{myers2014student, title = {Student Response Analysis}, author = {Myers, Sean and Parenti, Timothy and Thomas, Chris}, journal = {Technical Report - University of Pittsburgh Department of Computer Science}, year = {2014} }