The Evolution of the Web and Implications for an Incremental Crawler
Summary: Empirical four-month study of over 500K pages characterizes Web evolution for incremental crawling. Uses these dynamics to evaluate refresh policies and propose an architecture balancing index freshness, discovery of new pages, and crawl cost. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Junghoo Cho (Stanford University)
- 2. Hector Garcia-Molina (Stanford University)
BibTeX Citation
@article{cho_vldb00,
title = {{The Evolution of the Web and Implications for an Incremental Crawler}},
author = {Cho, Junghoo and Garcia-Molina, Hector},
journal = {PVLDB},
series = {{VLDB} '00},
pages = {200--209},
year = {2000}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 340 | Crawling the Hidden Web | 2001 | VLDB | 0.00020711119 |
| 8,556 | Effective Change Detection Using Sampling | 2002 | VLDB | 5.4119882e-05 |
| 12,527 | NEAR-Miner: Mining Evolution Associations of Web Site Directories for Efficient Maintenance of Web Archives | 2009 | VLDB | 5.093636e-05 |
| 13,726 | Dealing with Web Data: History and Look ahead | 2010 | VLDB | - |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,610 | Synchronizing a database to Improve Freshness | 2000 | SIGMOD | 0.00010220368 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,714 | Personalized PageRank on Evolving Graphs with an Incremental Index-Update Scheme | 2023 | SIGMOD |
| 2 | 340 | Crawling the Hidden Web | 2001 | VLDB |
| 3 | 12,426 | Optimizing Content Freshness of Relations Extracted From the Web Using Keyword Search | 2010 | SIGMOD |
| 4 | 9,683 | Optimal Algorithms for Crawling a Hidden Database in the Web | 2012 | VLDB |
| 5 | 8,705 | Progressive Deep Web Crawling Through Keyword Queries For Data Enrichment | 2019 | SIGMOD |
| 6 | 1,403 | Focused Crawling Using Context Graphs | 2000 | VLDB |
| 7 | 3,725 | Finding replicated web collections | 2000 | SIGMOD |
| 8 | 5,749 | RankMass Crawler: A Crawler with High Personalized PageRank Coverage Guarantee | 2007 | VLDB |
| 9 | 8,166 | Accurate and Efficient Crawling for Relevant Websites | 2004 | VLDB |
| 10 | 13,726 | Dealing with Web Data: History and Look ahead | 2010 | VLDB |