Passer à la navigation principale Passer à la recherche Passer au contenu principal

Automatic extraction of structured web data with domain knowledge

  • Nora Derouiche
  • , Bogdan Cautis
  • , Talel Abdessalem
  • CNRS LTCI

Résultats de recherche: Contribution à un journalArticle de conférenceRevue par des pairs

Résumé

We present in this paper a novel approach for extracting structured data from the Web, whose goal is to harvest real-world items from template-based HTML pages (the structured Web). It illustrates a two-phase querying of the Web, in which an intentional description of the data that is targeted is first provided, in a flexible and widely applicable manner. The extraction process leverages then both the input description and the source structure. Our approach is domain-independent, in the sense that it applies to any relation, either flat or nested, describing real-world items. Extensive experiments on five different domains and comparison with the main state of the art extraction systems from literature illustrate its flexibility and precision. We advocate via our technique that automatic extraction and integration of complex structured data can be done fast and effectively, when the redundancy of the Web meets knowledge over the to-be-extracted data.

langue originaleAnglais
Numéro d'article6228128
Pages (de - à)726-737
Nombre de pages12
journalProceedings - International Conference on Data Engineering
Les DOIs
étatPublié - 30 juil. 2012
Modification externeOui
EvénementIEEE 28th International Conference on Data Engineering, ICDE 2012 - Arlington, VA, États-Unis
Durée: 1 avr. 20125 avr. 2012

Empreinte digitale

Examiner les sujets de recherche de « Automatic extraction of structured web data with domain knowledge ». Ensemble, ils forment une empreinte digitale unique.

Contient cette citation