DBScholar

Back to papers

An Automatic Data Grabber for Large Web Sites

Summary: Automatically models large data-intensive websites intensionally as classes of structurally homogeneous pages with representative instances. It then infers one wrapper per class, using an external wrapper generator, enabling systematic site navigation and extraction. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
9285
Venue
VLDB
Year
2004
Pagerank
5.093636e-05
Overall Rank
12,783 | 12.30%
DOI
10.1016/B978-012088469-8.50137-6

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{crescenzi_vldb04,
        title = {{An Automatic Data Grabber for Large Web Sites}},
        author = {Crescenzi, Valter and Mecca, Giansalvatore and Merialdo, Paolo and Missier, Paolo},
        journal = {PVLDB},
        series = {{VLDB} '04},
        pages = {1321--1324},
        doi = {10.1016/B978-012088469-8.50137-6},
        url = {https://doi.org/10.1016/B978-012088469-8.50137-6},
        year = {2004}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 4 of 4 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
340 Crawling the Hidden Web 2001 VLDB 0.00020711119
503 RoadRunner: Towards Automatic Data Extraction from Large Web Sites 2001 VLDB 0.00017314037
588 Extracting Structured Data from Web Pages 2003 SIGMOD 0.00016092668
6,711 RoadRunner: Automatic Data Extraction from Data-Intensive Web Sites 2002 SIGMOD 5.7963597e-05
Previous Page 1 / 1 Next

Semantically Similar Papers