Passer à la navigation principale Passer à la recherche Passer au contenu principal

BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models

  • Zsolt T. Kardkovács
  • , Lynda Djennane
  • , Anna Field
  • , Boualem Benatallah
  • , Yacine Gaci
  • , Fabio Casati
  • , Walid Gaaloul
  • Dublin City University
  • École supérieure en Sciences et Technologies de l'Informatique et du Numérique
  • Plus que PRO Lab
  • ServiceNow
  • Università di Trento

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Sentiment Analysis (SA) models harbor inherent social biases that can be harmful in real-world applications. These biases are identified by examining the output of SA models for sentences that only vary in the identity groups of the subjects. Constructing natural, linguistically rich, relevant, and diverse sets of sentences that provide sufficient coverage over the domain is expensive, especially when addressing a wide range of biases: it requires domain experts and/or crowd-sourcing. In this paper, we present a novel bias testing framework, BTC-SAM, which generates high-quality test cases for bias testing in SA models with minimal specification using Large Language Models (LLMs) for the controllable generation of test sentences. Our experiments show that relying on LLMs can provide high linguistic variation and diversity in the test sentences, thereby offering better test coverage compared to base prompting methods even for previously unseen biases.

langue originaleAnglais
titreEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
rédacteurs en chefChristos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
EditeurAssociation for Computational Linguistics (ACL)
Pages15097-15113
Nombre de pages17
ISBN (Electronique)9798891763326
Les DOIs
étatPublié - 1 janv. 2025
Evénement30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 - Suzhou, Chine
Durée: 4 nov. 20259 nov. 2025

Série de publications

NomEMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference

Une conférence

Une conférence30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025
Pays/TerritoireChine
La villeSuzhou
période4/11/259/11/25

Empreinte digitale

Examiner les sujets de recherche de « BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation