Learning URI Selection Criteria to Improve the Crawling of Linked Open Data

Communication Dans Un Congrès Année : 2019

(1) , (1)

1 (France) 178918

CRISAM - Inria Sophia Antipolis - Méditerranée (2004 route des Lucioles BP 93 06902 Sophia Antipolis - France) 34586
- Inria - Institut National de Recherche en Informatique et en Automatique (Domaine de Voluceau Rocquencourt - BP 105 78153 Le Chesnay Cedex - France) 300009
Laboratoire I3S - SPARKS - Scalable and Pervasive softwARe and Knowledge Systems (Laboratoire I3S CS 40121 06903 Sophia Antipolis Cedex - France) 452156
- I3S - Laboratoire d'Informatique, Signaux, et Systèmes de Sophia Antipolis (2000, route des Lucioles - Les Algorithmes - bât. Euclide B 06900 Sophia Antipolis - France) 13009
  - UNS - Université Nice Sophia Antipolis (1965 - 2019) (Parc Valrose, 06100 Nice - France) 117617
  - CNRS - Centre National de la Recherche Scientifique : UMR7271 (France) 441569
  - UniCA - Université Côte d'Azur (Parc Valrose, 28, avenue Valrose 06108 Nice Cedex 2 - France) 1039632

"> WIMMICS - Web-Instrumented Man-Machine Interactions, Communities and Semantics

Hai Huang

Fonction : Auteur
PersonId : 1044348

Web-Instrumented Man-Machine Interactions, Communities and Semantics

Fabien Gandon

Fonction : Auteur
PersonId : 3342
IdHAL : fabien-gandon
ORCID : 0000-0003-0543-1232
IdRef : 076340074

Web-Instrumented Man-Machine Interactions, Communities and Semantics

Résumé

As the Web of Linked Open Data is growing the problem of crawling that cloud becomes increasingly important. Unlike normal Web crawlers, a Linked Data crawler performs a selection to focus on collecting linked RDF (including RDFa) data on the Web. From the perspectives of throughput and coverage, given a newly discovered and targeted URI, the key issue of Linked Data crawlers is to decide whether this URI is likely to dereference into an RDF data source and therefore it is worth downloading the representation it points to. Current solutions adopt heuristic rules to filter irrelevant URIs. Unfortunately, when the heuristics are too restrictive this hampers the coverage of crawling. In this paper, we propose and compare approaches to learn strategies for crawling Linked Data on the Web by predicting whether a newly discovered URI will lead to an RDF data source or not. We detail the features used in predicting the relevance and the methods we evaluated including a promising adaptation of FTRL-proximal online learning algorithm. We compare several options through extensive experiments including existing crawlers as baseline methods to evaluate their efficacy.

Mots clés

Linked Data Crawling Strategy Online Prediction Machine Learning

Domaines

Informatique [cs] Web Intelligence artificielle [cs.AI] Apprentissage [cs.LG]

Fichier principal

finalpp.pdf (380.74 Ko)

Origine	Fichiers produits par l'(les) auteur(s)

Hai Huang : Connectez-vous pour contacter le contributeur

https://inria.hal.science/hal-02073854

Soumis le : mercredi 20 mars 2019-11:53:41

Dernière modification le : lundi 26 février 2024-11:22:08

Archivage à long terme le : vendredi 21 juin 2019-13:56:31

Dates et versions

hal-02073854 , version 1 (20-03-2019)

Identifiants

HAL Id : hal-02073854 , version 1

Citer

Hai Huang, Fabien Gandon. Learning URI Selection Criteria to Improve the Crawling of Linked Open Data. ESWC2019 - 16th Extended Semantic Web Conference, Jun 2019, Portoroz, Slovenia. ⟨hal-02073854⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

UNIV-RENNES1 CNRS INRIA IRISA I3S WIMMICS INRIA2 UR1-MATH-STIC UR1-UFR-ISTIC UNIV-COTEDAZUR UNIV-RENNES 3IA-COTEDAZUR ANR UR1-MATH-NUM

625 Consultations

445 Téléchargements