| 694 |
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes |
2019 |
SIGMOD |
0.00014727089 |
| 972 |
The Data Civilizer System |
2017 |
CIDR |
0.00012763234 |
| 1,344 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00010956518 |
| 1,933 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
9.3459285e-05 |
| 1,957 |
SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search |
2024 |
SIGMOD |
9.3159589e-05 |
| 2,448 |
DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search |
2020 |
VLDB |
8.4494625e-05 |
| 3,350 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
7.3910669e-05 |
| 3,596 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
7.1790375e-05 |
| 4,493 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
6.5772891e-05 |
| 4,571 |
SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
6.5286445e-05 |
| 4,623 |
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach |
2016 |
SIGMOD |
6.4978811e-05 |
| 5,139 |
Efficient Load-Balanced Butterfly Counting on GPU |
2022 |
VLDB |
6.2579755e-05 |
| 5,549 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
6.088104e-05 |
| 5,700 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
6.0317753e-05 |
| 5,912 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
5.9501002e-05 |
| 5,920 |
ARKGraph: All-Range Approximate K-Nearest-Neighbor Graph |
2023 |
VLDB |
5.9475097e-05 |
| 5,953 |
Distributed Graph Simulation: Impossibility and Possibility |
2014 |
VLDB |
5.9359492e-05 |
| 6,605 |
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search |
2014 |
SIGMOD |
5.7367357e-05 |
| 6,768 |
Dynamic Range-Filtering Approximate Nearest Neighbor Search |
2025 |
VLDB |
5.686858e-05 |
| 7,030 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
5.6173605e-05 |
| 7,709 |
Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases |
2013 |
VLDB |
5.4717883e-05 |
| 7,833 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
5.443666e-05 |
| 7,906 |
Near-Duplicate Text Alignment with One Permutation Hashing |
2024 |
SIGMOD |
5.428873e-05 |
| 8,689 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
5.2905577e-05 |
| 9,126 |
Balance-Aware Distributed String Similarity-Based Query Processing System |
2019 |
VLDB |
5.22387e-05 |
| 9,879 |
META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion |
2016 |
VLDB |
5.1176637e-05 |
| 10,214 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
5.0596605e-05 |
| 10,478 |
LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH |
2026 |
SIGMOD |
4.9793485e-05 |
| 10,736 |
Near-Duplicate Text Alignment under Weighted Jaccard Similarity |
2026 |
VLDB |
4.9793485e-05 |
| 10,888 |
CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates |
2026 |
VLDB |
4.9793485e-05 |
| 11,850 |
SPINE: Scaling up Programming-by-Negative-Example for String Filtering and Transformation |
2022 |
SIGMOD |
4.9793485e-05 |