To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks
Summary: Cost-based optimizer for text-centric tasks chooses between crawl and query-based plans using a formal model of time and recall. Uses random-graph theory and statistics to estimate task-specific parameters; validated with large-scale experiments on three tasks and multiple real-life data sets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Panagiotis G. Ipeirotis (New York University)
- 2. Eugene Agichtein (Microsoft)
- 3. Pranay Jain (Columbia University)
- 4. Luis Gravano (Columbia University)
BibTeX Citation
@inproceedings{ipeirotis_sigmod06,
title = {{To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks}},
author = {Ipeirotis, Panagiotis G. and Agichtein, Eugene and Jain, Pranay and Gravano, Luis},
series = {{SIGMOD} '06},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1142473.1142504},
url = {https://dl.acm.org/doi/10.1145/1142473.1142504},
year = {2006}
}
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
| 508 | Random Sampling for Histogram Construction: How much is enough? | 1998 | SIGMOD | 0.00017275873 |
| 1,403 | Focused Crawling Using Context Graphs | 2000 | VLDB | 0.000108876 |
| 1,483 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB | 0.00010633832 |
| 1,601 | Automatic Discovery of Language Models for Text Databases | 1999 | SIGMOD | 0.00010239819 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 212 | Efficient IR-Style Keyword Search over Relational Databases | 2003 | VLDB |
| 2 | 2,948 | Optimized Query Execution in Large Search Engines with Global Page Ordering | 2003 | VLDB |
| 3 | 1,483 | Distributed Search over the Hidden Web: Hierarchical Database Sampling and Selection | 2002 | VLDB |
| 4 | 8,195 | When Speed Has a Price: Fast Information Extraction Using Approximate Algorithms | 2013 | VLDB |
| 5 | 1,403 | Focused Crawling Using Context Graphs | 2000 | VLDB |
| 6 | 12,287 | Probabilistic Query Rewriting for Efficient and Effective Keyword Search on Graph Data | 2013 | VLDB |
| 7 | 9,683 | Optimal Algorithms for Crawling a Hidden Database in the Web | 2012 | VLDB |
| 8 | 8,122 | A Formal Model of Trade-off between Optimization and Execution Costs in Semantic Query Optimization | 1988 | VLDB |
| 9 | 6,371 | QUEST: Query Optimization in Unstructured Document Analysis | 2025 | VLDB |
| 10 | 2,244 | Join Queries with External Text Sources: Execution and Optimization Techniques | 1995 | SIGMOD |