Passer à la navigation principale Passer à la recherche Passer au contenu principal

YAWN: A semantically annotated Wikipedia XML corpus

  • Max-Planck-Institut fur Informatik

Résultats de recherche: Le chapitre dans un livre, un rapport, une anthologie ou une collectionContribution à une conférenceRevue par des pairs

Résumé

The paper presents YAWN, a system to convert the well-known and widely used Wikipedia collection into an XML corpus with semantically rich, self-explaining tags. We introduce algorithms to annotate pages and links with concepts from the WordNet thesaurus. This annotation process exploits categorical information in Wikipedia, which is a high-quality, manually assigned source of information, extracts additional information from lists, and utilizes the invocations of templates with named parameters. We give examples how such annotations can be exploited for high-precision queries.

langue originaleAnglais
titreDatenbanksysteme in Business, Technologie und Web, BTW 2007 - 12th Fachtagung des GI-Fachbereichs "Datenbanken und Informationssysteme" (DBIS), Proceedings
Pages277-291
Nombre de pages15
étatPublié - 1 déc. 2007
Modification externeOui
Evénement12th Symposium of the German Informatics Society Section "Databases and Information Systems" (DBIS) on Database Systems in Business, Technology and Web, BTW 2007 - Aachen, Allemagne
Durée: 7 mars 20079 mars 2007

Série de publications

NomDatenbanksysteme in Business, Technologie und Web, BTW 2007 - 12th Fachtagung des GI-Fachbereichs "Datenbanken und Informationssysteme" (DBIS), Proceedings

Une conférence

Une conférence12th Symposium of the German Informatics Society Section "Databases and Information Systems" (DBIS) on Database Systems in Business, Technology and Web, BTW 2007
Pays/TerritoireAllemagne
La villeAachen
période7/03/079/03/07

Empreinte digitale

Examiner les sujets de recherche de « YAWN: A semantically annotated Wikipedia XML corpus ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation