Back to papers
RoadRunner: Towards Automatic Data Extraction from Large Web Sites
Summary: RoadRunner enables automatic data extraction from large web sites by generating wrappers via HTML page similarity/difference analysis. Real-world data-intensive site experiments demonstrate feasibility and scalability of the wrapper generation approach.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
hbdba100378cdcac0
Venue
VLDB
Year
2001
Pagerank
0.00016970049
Overall Rank
515 | 96.54%
DOI
-
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@article{crescenzi_vldb01,
title = {{RoadRunner: Towards Automatic Data Extraction from Large Web Sites}},
author = {Crescenzi, Valter and Mecca, Giansalvatore and Merialdo, Paolo},
journal = {PVLDB},
series = {{VLDB} '01},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 31 of 31 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
599
Extracting Structured Data from Web Pages
2003
SIGMOD
0.0001575536
622
On the Provenance of Non-Answers to Queries over Extracted Data
2008
VLDB
0.00015484312
1,337
Harvesting Relational Tables from Lists on the Web
2009
VLDB
0.00010987014
1,917
A Web of Concepts
2009
PODS
9.3790758e-05
2,120
An Analysis of Structured Data on the Web
2012
VLDB
9.0098024e-05
2,417
Instance-based Schema Matching for Web Databases by Domain-specific Query Probing
2004
VLDB
8.4930877e-05
2,599
Understanding Web Query Interfaces: Best-Effort Parsing with Hidden Syntax
2004
SIGMOD
8.2364103e-05
3,395
Automatic Wrappers for Large Scale Web Extraction
2011
VLDB
7.3436783e-05
3,568
Context-Aware Wrapping: Synchronized Data Extraction
2007
VLDB
7.200842e-05
3,588
Using the Structure of Web Sites for Automatic Segmentation of Tables
2004
SIGMOD
7.1870903e-05
3,759
Toward Large Scale Integration: Building a MetaQuerier over Databases on the Web
2005
CIDR
7.0430209e-05
3,912
Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets
2018
SIGMOD
6.9306061e-05
3,974
On the Complexity of Deriving Schema Mappings from Database Instances
2008
PODS
6.8868916e-05
4,010
Exploiting Content Redundancy for Web Information Extraction
2010
VLDB
6.8582009e-05
4,021
Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model
2009
SIGMOD
6.8498852e-05
5,319
Object-level Vertical Search
2007
CIDR
6.182759e-05
5,712
Documentum ECI Self-Repairing Wrappers: Performance Analysis
2006
SIGMOD
6.0218509e-05
6,015
WADaR: Joint Wrapper and Data Repair
2015
VLDB
5.9135138e-05
6,104
From Information to Knowledge: Harvesting Entities and Relationships from Web Sources
2010
PODS
5.886265e-05
6,398
CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web
2018
VLDB
5.7984902e-05
6,774
Optimal Schemes for Robust Web Extraction
2011
VLDB
5.6855924e-05
6,841
RoadRunner: Automatic Data Extraction from Data-Intensive Web Sites
2002
SIGMOD
5.6668827e-05
6,947
Web Data Extraction using Hybrid Program Synthesis: A Combination of Top-down and Bottom-up Inference
2020
SIGMOD
5.6369918e-05
7,579
Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents
2019
SIGMOD
5.4921926e-05
7,972
The Smallest Extraction Problem
2021
VLDB
5.4150414e-05
8,135
DEXTER: Large-Scale Discovery and Extraction of Product Specifications on the Web
2015
VLDB
5.3928122e-05
9,042
Measuring the Structural Similarity of Semistructured Documents Using Entropy
2007
VLDB
5.2318346e-05
12,744
ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data
2010
VLDB
4.9793485e-05
12,766
Building Ranked Mashups of Unstructured Sources with Uncertain Information
2010
VLDB
4.9793485e-05
13,008
Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
2006
VLDB
4.9793485e-05
13,073
An Automatic Data Grabber for Large Web Sites
2004
VLDB
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
2,374
Record-Boundary Discovery in Web Documents
1999
SIGMOD
2
599
Extracting Structured Data from Web Pages
2003
SIGMOD
3
4,021
Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model
2009
SIGMOD
4
8,725
An XML-based Wrapper Generator for Web Information Extraction
1999
SIGMOD
5
13,008
Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
2006
VLDB
6
6,774
Optimal Schemes for Robust Web Extraction
2011
VLDB
7
3,395
Automatic Wrappers for Large Scale Web Extraction
2011
VLDB
8
12,744
ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data
2010
VLDB
9
13,073
An Automatic Data Grabber for Large Web Sites
2004
VLDB
10
6,841
RoadRunner: Automatic Data Extraction from Data-Intensive Web Sites
2002
SIGMOD