An Efficient Filter for Approximate Membership Checking
Summary: Proposes a filter-verification framework for approximate substring membership against a large dictionary, enabling efficient named-entity and biomedical concept extraction. Introduces a novel in-memory filter that prunes non-matches early, guarantees no false negatives, and outperforms prior methods in filtering power and runtime. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Kaushik Chakrabarti (Microsoft)
- 2. Surajit Chaudhuri (Microsoft)
- 3. Venkatesh Ganti (Microsoft)
- 4. Dong Xin (Microsoft)
BibTeX Citation
@inproceedings{chakrabarti_sigmod08,
title = {{An Efficient Filter for Approximate Membership Checking}},
author = {Chakrabarti, Kaushik and Chaudhuri, Surajit and Ganti, Venkatesh and Xin, Dong},
series = {{SIGMOD} '08},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1376616.1376697},
url = {https://dl.acm.org/doi/10.1145/1376616.1376697},
year = {2008}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5 | Optimal Aggregation Algorithms for Middleware [Extended Abstract] | 2001 | PODS | 0.0010828372 |
| 21 | Similarity Search in High Dimensions via Hashing | 1999 | VLDB | 0.00056760516 |
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,504 | TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching | 2023 | VLDB |
| 2 | 4,396 | Approximate String Joins with Abbreviations | 2018 | VLDB |
| 3 | 5,412 | Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction | 2011 | SIGMOD |
| 4 | 10,083 | ChainedFilter: Combining Membership Filters by Chain Rule | 2023 | SIGMOD |
| 5 | 10,086 | Efficient and Effective KNN Sequence Search with Approximate n-grams | 2014 | VLDB |
| 6 | 6,484 | A Pivotal Prefix Based Filtering Algorithm for String Similarity Search | 2014 | SIGMOD |
| 7 | 8,114 | Approximate Substring Matching over Uncertain Strings | 2011 | VLDB |
| 8 | 3,446 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD |
| 9 | 10,085 | Local Filtering: Improving the Performance of Approximate Queries on String Collections | 2015 | SIGMOD |
| 10 | 7,692 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |