Passer à la navigation principale Passer à la recherche Passer au contenu principal

Adding missing words to regular expressions

  • Telecom Paris
  • 95014 Cergy

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Regular expressions (regexes) are patterns that are used in many applications to extract words or tokens from text. However, even hand-crafted regexes may fail to match all the intended words. In this paper, we propose a novel way to generalize a given regex so that it matches also a set of missing (previously non-matched) words. Our method finds an approximate match between the missing words and the regex, and adds disjunctions for the unmatched parts appropriately. We show that this method can not just improve the precision and recall of the regex, but also generate much shorter regexes than baselines and competitors on various datasets.

langue originaleAnglais
titreAdvances in Knowledge Discovery and Data Mining - 22nd Pacific-Asia Conference, PAKDD 2018, Proceedings
rédacteurs en chefDinh Phung, Vincent S. Tseng, Geoffrey I. Webb, Bao Ho, Mohadeseh Ganji, Lida Rashidi
EditeurSpringer Verlag
Pages67-79
Nombre de pages13
ISBN (imprimé)9783319930367
Les DOIs
étatPublié - 1 janv. 2018
Evénement22nd Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, PAKDD 2018 - Melbourne, Australie
Durée: 3 juin 20186 juin 2018

Série de publications

NomLecture Notes in Computer Science
Volume10938 LNAI
ISSN (imprimé)0302-9743
ISSN (Electronique)1611-3349

Une conférence

Une conférence22nd Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, PAKDD 2018
Pays/TerritoireAustralie
La villeMelbourne
période3/06/186/06/18

Empreinte digitale

Examiner les sujets de recherche de « Adding missing words to regular expressions ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation