TY - GEN
T1 - Empirical Evaluation of Social Bias in Text Classification Systems
AU - Djennane, Lynda
AU - Kardkovacs, Zsolt T.
AU - Benatallah, Boualem
AU - Gaci, Yacine
AU - Gaaloul, Walid
AU - Farah, Zoubeyr
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024/1/1
Y1 - 2024/1/1
N2 - Social bias stereotypes have recently raised significant ethical concerns in Natural Language Processing (NLP). NLP models, particularly those used for text classification, often perpetuate these biases by producing different output scores for various demographic groups, leading to discriminatory outcomes. In this paper, we conduct a comprehensive evaluation of potential social biases in a diverse array of text classification tasks, focusing on gender, race, and religion through counterfactual fairness testing. We examined 11 widely-used text classification models from Hugging Face and 3 commercial sentiment analysis models using 5 different datasets. Our findings reveal a pronounced tendency for these systems to favour certain demographic groups over others, with statistically significant biases detected. Specifically, the analysis highlights substantial disparities in how these models score identical content when demographic variables are altered, demonstrating inherent biases in the underlying models.
AB - Social bias stereotypes have recently raised significant ethical concerns in Natural Language Processing (NLP). NLP models, particularly those used for text classification, often perpetuate these biases by producing different output scores for various demographic groups, leading to discriminatory outcomes. In this paper, we conduct a comprehensive evaluation of potential social biases in a diverse array of text classification tasks, focusing on gender, race, and religion through counterfactual fairness testing. We examined 11 widely-used text classification models from Hugging Face and 3 commercial sentiment analysis models using 5 different datasets. Our findings reveal a pronounced tendency for these systems to favour certain demographic groups over others, with statistically significant biases detected. Specifically, the analysis highlights substantial disparities in how these models score identical content when demographic variables are altered, demonstrating inherent biases in the underlying models.
KW - Natural language processing
KW - social bias
KW - text classification models
UR - https://www.scopus.com/pages/publications/85218342364
U2 - 10.1109/FLLM63129.2024.10852435
DO - 10.1109/FLLM63129.2024.10852435
M3 - Conference contribution
AN - SCOPUS:85218342364
T3 - 2024 2nd International Conference on Foundation and Large Language Models, FLLM 2024
SP - 313
EP - 321
BT - 2024 2nd International Conference on Foundation and Large Language Models, FLLM 2024
A2 - Jararweh, Yaser
A2 - Jansen, Jim
A2 - Alsmirat, Mohammad
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2nd International Conference on Foundation and Large Language Models, FLLM 2024
Y2 - 26 November 2024 through 29 November 2024
ER -