Passer à la navigation principale Passer à la recherche Passer au contenu principal

Character and subword-based word representation for neural language modeling prediction

  • CNRS

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

16 Citations (Scopus)

Résumé

Most of neural language models use different kinds of embeddings for word prediction. While word embeddings can be associated to each word in the vocabulary or derived from characters as well as factored morphological decomposition, these word representations are mainly used to parametrize the input, i.e. the context of prediction. This work investigates the effect of using subword units (character and factored morphological decomposition) to build output representations for neural language modeling. We present a case study on Czech, a morphologically-rich language, experimenting with different input and output representations. When working with the full training vocabulary, despite unstable training, our experiments show that augmenting the output word representations with character-based embeddings can significantly improve the performance of the model. Moreover, reducing the size of the output look-up table, to let the character-based embeddings represent rare words, brings further improvement.

langue originaleAnglais
titreEMNLP 2017 - 1st Workshop on Subword and Character Level Models in NLP, SCLeM 2017 - Proceedings of the Workshop
rédacteurs en chefManaal Faruqui, Hinrich Schutze, Isabel Trancoso, Yaghoobzadeh Yadollah
EditeurAssociation for Computational Linguistics (ACL)
Pages1-13
Nombre de pages13
ISBN (Electronique)9781945626913
étatPublié - 1 janv. 2017
EvénementEMNLP 2017 1st Workshop on Subword and Character Level Models in NLP, SCLeM 2017 - Copenhagen, Danemark
Durée: 7 sept. 2017 → …

Série de publications

NomEMNLP 2017 - 1st Workshop on Subword and Character Level Models in NLP, SCLeM 2017 - Proceedings of the Workshop

Une conférence

Une conférenceEMNLP 2017 1st Workshop on Subword and Character Level Models in NLP, SCLeM 2017
Pays/TerritoireDanemark
La villeCopenhagen
période7/09/17 → …

Empreinte digitale

Examiner les sujets de recherche de « Character and subword-based word representation for neural language modeling prediction ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation