Benchmarking Declarative Approximate Selection Predicates
Summary: Benchmarks declarative approximate selection predicates; introduces probabilistic similarity predicates for data quality using language models and HMMs with declarative realization. Classifies existing predicates by class and reports runtime and accuracy for data-cleaning tasks. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Amit Chandel (University of Toronto)
- 2. Oktie Hassanzadeh (University of Toronto)
- 3. Nick Koudas (University of Toronto)
- 4. Mohammad Sadoghi (University of Toronto)
- 5. Divesh Srivastava (AT&T)
BibTeX Citation
@inproceedings{chandel_sigmod07,
title = {{Benchmarking Declarative Approximate Selection Predicates}},
author = {Chandel, Amit and Hassanzadeh, Oktie and Koudas, Nick and Sadoghi, Mohammad and Srivastava, Divesh},
series = {{SIGMOD} '07},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1247480.1247521},
url = {https://dl.acm.org/doi/10.1145/1247480.1247521},
year = {2007}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
| 160 | Integration of Heterogeneous Databases Without Common Domains Using Queries Based on Textual Similarity | 1998 | SIGMOD | 0.0002802209 |
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
| 201 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB | 0.00025558602 |
| 254 | Record Linkage: Similarity Measures and Algorithms | 2006 | SIGMOD | 0.00023199211 |
| 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB | 0.00021839661 |
| 4,009 | Flexible String Matching Against Large Databases in Practice | 2004 | VLDB | 6.9612537e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,162 | Efficiently Approximating Selectivity Functions using Low Overhead Regression Models | 2020 | VLDB |
| 2 | 280 | Selectivity Estimation using Probabilistic Models | 2001 | SIGMOD |
| 3 | 4,055 | Hybrid In-Database Inference for Declarative Information Extraction | 2011 | SIGMOD |
| 4 | 6,674 | Benchmarking Approximate Consistent Query Answering | 2021 | PODS |
| 5 | 3,610 | Merging the Results of Approximate Match Operations | 2004 | VLDB |
| 6 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 7 | 50 | Efficient Query Evaluation on Probabilistic Databases | 2004 | VLDB |
| 8 | 1,573 | Deep Learning Models for Selectivity Estimation of Multi-Attribute Queries | 2020 | SIGMOD |
| 9 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 10 | 4,136 | Approximating Predicates and Expressive Queries on Probabilistic Databases | 2008 | PODS |