Back to papers
RoadRunner: Towards Automatic Data Extraction from Large Web Sites
Summary: RoadRunner enables automatic data extraction from large web sites by generating wrappers via HTML page similarity/difference analysis. Real-world data-intensive site experiments demonstrate feasibility and scalability of the wrapper generation approach.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
8926
Venue
VLDB
Year
2001
Pagerank
0.00017314037
Overall Rank
503 | 96.56%
DOI
-
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@article{crescenzi_vldb01,
title = {{RoadRunner: Towards Automatic Data Extraction from Large Web Sites}},
author = {Crescenzi, Valter and Mecca, Giansalvatore and Merialdo, Paolo},
journal = {PVLDB},
series = {{VLDB} '01},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 31 of 31 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
588
Extracting Structured Data from Web Pages
2003
SIGMOD
0.00016092668
628
On the Provenance of Non-Answers to Queries over Extracted Data
2008
VLDB
0.00015630285
1,316
Harvesting Relational Tables from Lists on the Web
2009
VLDB
0.00011181216
1,873
A Web of Concepts
2009
PODS
9.5751398e-05
2,096
An Analysis of Structured Data on the Web
2012
VLDB
9.1746663e-05
2,372
Instance-based Schema Matching for Web Databases by Domain-specific Query Probing
2004
VLDB
8.6786716e-05
2,554
Understanding Web Query Interfaces: Best-Effort Parsing with Hidden Syntax
2004
SIGMOD
8.4252755e-05
3,340
Automatic Wrappers for Large Scale Web Extraction
2011
VLDB
7.5040045e-05
3,505
Context-Aware Wrapping: Synchronized Data Extraction
2007
VLDB
7.3582812e-05
3,523
Using the Structure of Web Sites for Automatic Segmentation of Tables
2004
SIGMOD
7.3452966e-05
3,690
Toward Large Scale Integration: Building a MetaQuerier over Databases on the Web
2005
CIDR
7.2007927e-05
3,894
Navigating the Data Lake with DATAMARAN: Automatically Extracting Structure from Log Datasets
2018
SIGMOD
7.0417782e-05
3,942
On the Complexity of Deriving Schema Mappings from Database Instances
2008
PODS
7.0063604e-05
3,943
Exploiting Content Redundancy for Web Information Extraction
2010
VLDB
7.0063237e-05
3,944
Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model
2009
SIGMOD
7.0058157e-05
5,235
Object-level Vertical Search
2007
CIDR
6.3047329e-05
5,585
Documentum ECI Self-Repairing Wrappers: Performance Analysis
2006
SIGMOD
6.1599116e-05
5,915
WADaR: Joint Wrapper and Data Repair
2015
VLDB
6.0437843e-05
6,011
From Information to Knowledge: Harvesting Entities and Relationships from Web Sources
2010
PODS
6.0105397e-05
6,320
CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web
2018
VLDB
5.9150411e-05
6,644
Optimal Schemes for Robust Web Extraction
2011
VLDB
5.8150539e-05
6,711
RoadRunner: Automatic Data Extraction from Data-Intensive Web Sites
2002
SIGMOD
5.7963597e-05
6,809
Web Data Extraction using Hybrid Program Synthesis: A Combination of Top-down and Bottom-up Inference
2020
SIGMOD
5.7663737e-05
7,811
The Smallest Extraction Problem
2021
VLDB
5.5393291e-05
7,969
DEXTER: Large-Scale Discovery and Extraction of Product Specifications on the Web
2015
VLDB
5.5165897e-05
8,677
Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents
2019
SIGMOD
5.3867222e-05
8,882
Measuring the Structural Similarity of Semistructured Documents Using Entropy
2007
VLDB
5.3516704e-05
12,453
ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data
2010
VLDB
5.093636e-05
12,475
Building Ranked Mashups of Unstructured Sources with Uncertain Information
2010
VLDB
5.093636e-05
12,718
Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
2006
VLDB
5.093636e-05
12,783
An Automatic Data Grabber for Large Web Sites
2004
VLDB
5.093636e-05
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
2,329
Record-Boundary Discovery in Web Documents
1999
SIGMOD
2
588
Extracting Structured Data from Web Pages
2003
SIGMOD
3
3,944
Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model
2009
SIGMOD
4
8,558
An XML-based Wrapper Generator for Web Information Extraction
1999
SIGMOD
5
12,718
Automatic Extraction of Dynamic Record Sections From Search Engine Result Pages
2006
VLDB
6
6,644
Optimal Schemes for Robust Web Extraction
2011
VLDB
7
3,340
Automatic Wrappers for Large Scale Web Extraction
2011
VLDB
8
12,453
ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data
2010
VLDB
9
12,783
An Automatic Data Grabber for Large Web Sites
2004
VLDB
10
6,711
RoadRunner: Automatic Data Extraction from Data-Intensive Web Sites
2002
SIGMOD