Probe, Count, and Classify: Categorizing Hidden-Web Databases
Summary: Automates hidden-web database categorization with a small set of query probes; uses per-probe match counts, no page retrieval. Evaluated on 100+ real databases; achieves low overhead and high accuracy for automatic hierarchical categorization. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Panagiotis G. Ipeirotis (Columbia University)
- 2. Luis Gravano (Columbia University)
- 3. Mehran Sahami (Epiphany Inc.)
BibTeX Citation
@inproceedings{ipeirotis_sigmod01,
title = {{Probe, Count, and Classify: Categorizing Hidden-Web Databases}},
author = {Ipeirotis, Panagiotis G. and Gravano, Luis and Sahami, Mehran},
series = {{SIGMOD} '01},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/375663.375671},
url = {https://dl.acm.org/doi/10.1145/375663.375671},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,517 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB | 0.00010400233 |
| 2,913 | Knocking the Door to the Deep Web: Integrating Web Query Interfaces | 2004 | SIGMOD | 7.8643473e-05 |
| 5,985 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD | 5.926154e-05 |
| 13,117 | From Focused Crawling to Expert Information: an Application Framework for Web Exploration and Portal Generation | 2003 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,631 | Automatic Discovery of Language Models for Text Databases | 1999 | SIGMOD | 0.00010022634 |
| 1,655 | Determining Text Databases to Search in the Internet | 1998 | VLDB | 9.9733472e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,985 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD |
| 2 | 2,417 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing | 2004 | VLDB |
| 3 | 89 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB |
| 4 | 3,346 | Distributed Hypertext Resource Discovery Through Examples | 1999 | VLDB |
| 5 | 1,582 | Google’s Deep-Web Crawl | 2008 | VLDB |
| 6 | 8,736 | Progressive Deep Web Crawling Through Keyword Queries For Data Enrichment | 2019 | SIGMOD |
| 7 | 348 | Crawling the Hidden Web | 2001 | VLDB |
| 8 | 8,833 | Unbiased Estimation of Size and Other Aggregates Over Hidden Web Databases | 2010 | SIGMOD |
| 9 | 1,517 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB |
| 10 | 9,856 | Optimal Algorithms for Crawling a Hidden Database in the Web | 2012 | VLDB |