Passer à la navigation principale Passer à la recherche Passer au contenu principal

The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations

  • Institut Polytechnique de Paris
  • INRIA Institut National de Recherche en Informatique et en Automatique

Résultats de recherche: Contribution à un journalArticleRevue par des pairs

2 Citations (Scopus)

Résumé

When deriving contextualized word representations from language models, a decision needs to be made on how to obtain one for out-of-vocabulary (OOV) words that are segmented into subwords. What is the best way to represent these words with a single vector, and are these representations of worse quality than those of in-vocabulary words? We carry out an intrinsic evaluation of embeddings from different models on semantic similarity tasks involving OOV words. Our analysis reveals, among other interesting findings, that the quality of representations of words that are split is often, but not always, worse than that of the embeddings of known words. Their similarity values, however, must be interpreted with caution.

langue originaleAnglais
Pages (de - à)299-320
Nombre de pages22
journalTransactions of the Association for Computational Linguistics
Volume12
Les DOIs
étatPublié - 1 janv. 2024

Empreinte digitale

Examiner les sujets de recherche de « The Impact of Word Splitting on the Semantic Content of Contextualized Word Representations ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation