Efficient set joins on similarity predicates
Summary: General, scalable algorithm for set joins on similarity predicates (intersect size, Jaccard, cosine, edit distance) extending beyond simple containment. Inverted-index probing with staged optimizations, memory-efficient partitioning, and index compression enabling in-memory operation; generalizes to weighted/unweighted partial word overlap. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sunita Sarawagi (Indian Institute of Technology Mumbai)
- 2. Alok Kirpal (Indian Institute of Technology Mumbai; Yahoo)
BibTeX Citation
@inproceedings{sarawagi_sigmod04,
title = {{Efficient set joins on similarity predicates}},
author = {Sarawagi, Sunita and Kirpal, Alok},
series = {{SIGMOD} '04},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1007568.1007652},
url = {https://dl.acm.org/doi/10.1145/1007568.1007652},
year = {2004}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 53 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 27 | Fast Algorithms for Mining Association Rules | 1994 | VLDB | 0.00052255472 |
| 107 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033511706 |
| 158 | Robust and Efficient Fuzzy Match for Online Data Cleaning | 2003 | SIGMOD | 0.00028199923 |
| 210 | An Evaluation of Non-Equijoin Algorithms | 1991 | VLDB | 0.00024797689 |
| 306 | Eliminating Fuzzy Duplicates in Data Warehouses | 2002 | VLDB | 0.00021839661 |
| 1,151 | Set Containment Joins: The Good, The Bad and The Ugly | 2000 | VLDB | 0.00011938186 |
| 1,568 | Evaluation of Main Memory Join Algorithms for Joins with Subset Join Predicates | 1997 | VLDB | 0.0001034191 |
| 1,592 | Efficient Processing of Joins on Set-valued Attributes | 2003 | SIGMOD | 0.00010253457 |
| 2,061 | Selectivity Estimation For Boolean Queries | 2000 | PODS | 9.2445049e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,951 | Extensible and Robust Evaluation of Similarity Queries | 2025 | VLDB |
| 2 | 356 | Efficient Parallel Set-Similarity Joins Using MapReduce | 2010 | SIGMOD |
| 3 | 9,052 | Fast Approximate Similarity Join in Vector Databases | 2025 | SIGMOD |
| 4 | 6,906 | Efficient Similarity Join and Search on Multi-Attribute Data | 2015 | SIGMOD |
| 5 | 3,724 | Overlap Set Similarity Joins with Theoretical Guarantees | 2018 | SIGMOD |
| 6 | 3,474 | An Efficient Partition Based Method for Exact Set Similarity Joins | 2016 | VLDB |
| 7 | 1,592 | Efficient Processing of Joins on Set-valued Attributes | 2003 | SIGMOD |
| 8 | 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB |
| 9 | 2,501 | An Empirical Evaluation of Set Similarity Join Techniques | 2016 | VLDB |
| 10 | 3,040 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB |