Passer à la navigation principale Passer à la recherche Passer au contenu principal

Quantifying the Bias of Transformer-Based Language Models for African American English in Masked Language Modeling

  • University College London
  • Telecom Paris

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

6 Citations (Scopus)

Résumé

In recent years, groundbreaking transformer-based language models (LMs) have made tremendous advances in natural language processing (NLP) tasks. However, the measurement of their fairness with respect to different social groups still remains unsolved. In this paper, we propose and thoroughly validate an evaluation technique to assess the quality and bias of language model predictions on transcripts of both spoken African American English (AAE) and Spoken American English (SAE). Our analysis reveals the presence of a bias towards SAE encoded by state-of-the-art LMs such as BERT and DistilBERT and a lower bias in distilled LMs. We also observe a bias towards AAE in RoBERTa and BART. Additionally, we show evidence that this disparity is present across all the LMs when we only consider the grammar and the syntax specific to AAE.

langue originaleAnglais
titreAdvances in Knowledge Discovery and Data Mining - 27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023, Proceedings
rédacteurs en chefHisashi Kashima, Tsuyoshi Ide, Wen-Chih Peng
EditeurSpringer Science and Business Media Deutschland GmbH
Pages532-543
Nombre de pages12
ISBN (imprimé)9783031333736
Les DOIs
étatPublié - 1 janv. 2023
Evénement27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023 - Hybrid, Osaka, Japon
Durée: 25 mai 202328 mai 2023

Série de publications

NomLecture Notes in Computer Science
Volume13935 LNCS
ISSN (imprimé)0302-9743
ISSN (Electronique)1611-3349

Une conférence

Une conférence27th Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2023
Pays/TerritoireJapon
La villeHybrid, Osaka
période25/05/2328/05/23

Empreinte digitale

Examiner les sujets de recherche de « Quantifying the Bias of Transformer-Based Language Models for African American English in Masked Language Modeling ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation