Probe, Count, and Classify: Categorizing Hidden-Web Databases
Summary: Automates hidden-web database categorization with a small set of query probes; uses per-probe match counts, no page retrieval. Evaluated on 100+ real databases; achieves low overhead and high accuracy for automatic hierarchical categorization. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Panagiotis G. Ipeirotis (Columbia University)
- 2. Luis Gravano (Columbia University)
- 3. Mehran Sahami (Epiphany Inc.)
BibTeX Citation
@inproceedings{ipeirotis_sigmod01,
title = {{Probe, Count, and Classify: Categorizing Hidden-Web Databases}},
author = {Ipeirotis, Panagiotis G. and Gravano, Luis and Sahami, Mehran},
series = {{SIGMOD} '01},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/375663.375671},
url = {https://dl.acm.org/doi/10.1145/375663.375671},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,517 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB | 0.00010395336 |
| 2,914 | Knocking the Door to the Deep Web: Integrating Web Query Interfaces | 2004 | SIGMOD | 7.8607738e-05 |
| 5,985 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD | 5.9233486e-05 |
| 13,123 | From Focused Crawling to Expert Information: an Application Framework for Web Exploration and Portal Generation | 2003 | VLDB | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,632 | Automatic Discovery of Language Models for Text Databases | 1999 | SIGMOD | 0.00010017983 |
| 1,656 | Determining Text Databases to Search in the Internet | 1998 | VLDB | 9.9686565e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,985 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD |
| 2 | 2,419 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing | 2004 | VLDB |
| 3 | 89 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB |
| 4 | 3,347 | Distributed Hypertext Resource Discovery Through Examples | 1999 | VLDB |
| 5 | 1,582 | Google’s Deep-Web Crawl | 2008 | VLDB |
| 6 | 8,744 | Progressive Deep Web Crawling Through Keyword Queries For Data Enrichment | 2019 | SIGMOD |
| 7 | 348 | Crawling the Hidden Web | 2001 | VLDB |
| 8 | 8,842 | Unbiased Estimation of Size and Other Aggregates Over Hidden Web Databases | 2010 | SIGMOD |
| 9 | 1,517 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB |
| 10 | 9,863 | Optimal Algorithms for Crawling a Hidden Database in the Web | 2012 | VLDB |