Passer à la navigation principale Passer à la recherche Passer au contenu principal

Enhancing bias detection in text classification models through high-quality diverse test cases

  • Lynda Djennane
  • , Zsolt T. Kardkovács
  • , Boualem Benatallah
  • , Yacine Gaci
  • , Walid Gaaloul
  • , Zoubeyr Farah
  • École supérieure en Sciences et Technologies de l’Informatique et du Numérique
  • Dublin City University
  • Plus que PRO Lab

Résultats de recherche: Contribution à un journalArticleRevue par des pairs

Résumé

Context: Evaluating social biases in text classification systems is essential for developing fair NLP technologies. However, most existing approaches rely on fixed, hand-crafted templates that provide limited linguistic coverage and often fail to expose subtle or context-dependent biased behaviors in modern language models. Objective: This work investigates whether introducing linguistic diversity—through controlled paraphrasing—can uncover latent social biases that remain undetected under traditional template-based evaluations. Methods: We introduce a pipeline that generates high-quality, meaning-preserving paraphrases varying in lexical and syntactic structure. These diverse test cases are applied across 17 widely used text classification models spanning seven tasks and nine social bias categories. Bias is quantified before and after introducing paraphrases to measure the additional biased behavior revealed. Results: Diverse test cases substantially increase bias detection coverage. Paraphrased sentences reveal between 6% and 65% additional biased predictions depending on the task. Linguistic transformations such as reordering, synonym substitution, and question-style phrasing are particularly effective at exposing latent biases, especially for categories such as nationality, race, and family status. Several models that appeared unbiased under fixed templates exhibited biased behavior once linguistic variation was introduced. Conclusion: Linguistic diversity plays a crucial role in uncovering subtle and context-dependent biases in text classification. The proposed pipeline demonstrates that template-based methods alone are insufficient for reliable fairness evaluation and that diverse paraphrased test cases provide a practical and effective means to expand bias coverage. This work offers methodological guidance and test resources for more robust and meaningful bias assessment in NLP systems.

langue originaleAnglais
Numéro d'article108253
journalInformation and Software Technology
Volume199
Les DOIs
étatPublié - 1 nov. 2026

Empreinte digitale

Examiner les sujets de recherche de « Enhancing bias detection in text classification models through high-quality diverse test cases ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation