Passer à la navigation principale Passer à la recherche Passer au contenu principal

Canonicalizing open knowledge bases

  • Telecom Paris
  • Google Inc.

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

Open information extraction approaches have led to the creation of large knowledge bases from the Web. The problem with such methods is that their entities and relations are not canonicalized, leading to redundant and ambiguous facts. For example, they may store (Barack Obama, was born in, Honolulu) and (Obama, place of birth, Honolulu). In this paper, we present an approach based on machine learning methods that can canonicalize such Open IE triples, by clustering synonymous names and phrases. We also provide a detailed discussion about the different signals, features and design choices that influence the quality of synonym resolution for noun phrases in Open IE KBs, thus shedding light on the middle ground between "open" and "closed" information extraction systems.

langue originaleAnglais
titreCIKM 2014 - Proceedings of the 2014 ACM International Conference on Information and Knowledge Management
EditeurAssociation for Computing Machinery
Pages1679-1688
Nombre de pages10
ISBN (Electronique)9781450325981
Les DOIs
étatPublié - 3 nov. 2014
Evénement23rd ACM International Conference on Information and Knowledge Management, CIKM 2014 - Shanghai, Chine
Durée: 3 nov. 20147 nov. 2014

Série de publications

NomCIKM 2014 - Proceedings of the 2014 ACM International Conference on Information and Knowledge Management

Une conférence

Une conférence23rd ACM International Conference on Information and Knowledge Management, CIKM 2014
Pays/TerritoireChine
La villeShanghai
période3/11/147/11/14

Empreinte digitale

Examiner les sujets de recherche de « Canonicalizing open knowledge bases ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation