Skip to main navigation Skip to search Skip to main content

Characterizing Cross-Contamination on Multitask Multimodal Large Language Models

  • Universidad Carlos III de Madrid
  • Telecom Sudparis
  • Université Paris-Saclay

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advances in Large Language Models (LLMs) have paved the way for the emergence of Multimodal and Multitask LLMs (MMLLMs). While poisoning attacks are well studied in traditional AI models, they have not been characterized for MMLLMs. This paper focuses on a MMLLM-specific threat, dubbed cross-contamination, by which one task may be indirectly affected by the poison inserted in other tasks. In this vein, this paper addresses representative data poisoning techniques (label flipping and backdooring) across different text-, image- and video-based tasks. We analyse the effect on three representative models, namely OFA, LLaMa Vision 3.2 and Qwen2 Visual. The size of the dataset, the poison budget, the number of epochs and the generation strategy are analysed. Results confirm that poisoning can spread through unrelated tasks while preserving the overall model performance. Interestingly, attack effectivity rates above 99% can be achieved with just 5% of poison. The level of stealthiness depends on the type of attack, the targeted task and the model. We have also found a surprise effect by which a higher number of epochs do not always lead to a greater spread of the poison. Lastly, the choice of generation strategies has led to differences beyond 70% in the attack success rate.

Original languageEnglish
JournalInformation Systems Frontiers
DOIs
Publication statusAccepted/In press - 1 Jan 2025

Keywords

  • Backdooring
  • Cross-contamination
  • Data poisoning
  • Machine learning
  • Multimodal LLMs
  • Multitasking LLMs

Fingerprint

Dive into the research topics of 'Characterizing Cross-Contamination on Multitask Multimodal Large Language Models'. Together they form a unique fingerprint.

Cite this