Probe, Count, and Classify: Categorizing Hidden-Web Databases
Summary: Automates hidden-web database categorization with a small set of query probes; uses per-probe match counts, no page retrieval. Evaluated on 100+ real databases; achieves low overhead and high accuracy for automatic hierarchical categorization. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Panagiotis G. Ipeirotis (Columbia University)
- 2. Luis Gravano (Columbia University)
- 3. Mehran Sahami (Epiphany Inc.)
BibTeX Citation
@inproceedings{ipeirotis_sigmod01,
title = {{Probe, Count, and Classify: Categorizing Hidden-Web Databases}},
author = {Ipeirotis, Panagiotis G. and Gravano, Luis and Sahami, Mehran},
series = {{SIGMOD} '01},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/375663.375671},
url = {https://dl.acm.org/doi/10.1145/375663.375671},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,483 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB | 0.00010633832 |
| 2,916 | Knocking the Door to the Deep Web: Integrating Web Query Interfaces | 2004 | SIGMOD | 7.9664495e-05 |
| 5,868 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD | 6.0613238e-05 |
| 12,827 | From Focused Crawling to Expert Information: an Application Framework for Web Exploration and Portal Generation | 2003 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,601 | Automatic Discovery of Language Models for Text Databases | 1999 | SIGMOD | 0.00010239819 |
| 1,625 | Determining Text Databases to Search in the Internet | 1998 | VLDB | 0.00010191388 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,868 | A Random Walk Approach to Sampling Hidden Databases | 2007 | SIGMOD |
| 2 | 2,372 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing | 2004 | VLDB |
| 3 | 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB |
| 4 | 3,383 | Distributed Hypertext Resource Discovery Through Examples | 1999 | VLDB |
| 5 | 1,554 | Google’s Deep-Web Crawl | 2008 | VLDB |
| 6 | 8,705 | Progressive Deep Web Crawling Through Keyword Queries For Data Enrichment | 2019 | SIGMOD |
| 7 | 340 | Crawling the Hidden Web | 2001 | VLDB |
| 8 | 8,671 | Unbiased Estimation of Size and Other Aggregates Over Hidden Web Databases | 2010 | SIGMOD |
| 9 | 1,483 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB |
| 10 | 9,683 | Optimal Algorithms for Crawling a Hidden Database in the Web | 2012 | VLDB |