| 694 |
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes |
2019 |
SIGMOD |
58 |
0.00014721161 |
| 972 |
The Data Civilizer System |
2017 |
CIDR |
56 |
0.00012757732 |
| 1,344 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
50 |
0.00010951939 |
| 1,934 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
39 |
9.341845e-05 |
| 1,958 |
SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search |
2024 |
SIGMOD |
37 |
9.3118048e-05 |
| 2,444 |
DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search |
2020 |
VLDB |
19 |
8.4556406e-05 |
| 3,350 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
22 |
7.3875743e-05 |
| 3,598 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
17 |
7.1756405e-05 |
| 4,496 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
6 |
6.5741786e-05 |
| 4,573 |
SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
11 |
6.5257221e-05 |
| 4,625 |
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach |
2016 |
SIGMOD |
26 |
6.4948389e-05 |
| 5,141 |
Efficient Load-Balanced Butterfly Counting on GPU |
2022 |
VLDB |
6 |
6.2550131e-05 |
| 5,551 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
11 |
6.0852227e-05 |
| 5,703 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
7 |
6.02892e-05 |
| 5,903 |
ARKGraph: All-Range Approximate K-Nearest-Neighbor Graph |
2023 |
VLDB |
11 |
5.9512758e-05 |
| 5,914 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
9 |
5.9472869e-05 |
| 5,955 |
Distributed Graph Simulation: Impossibility and Possibility |
2014 |
VLDB |
11 |
5.9331392e-05 |
| 6,132 |
Dynamic Range-Filtering Approximate Nearest Neighbor Search |
2025 |
VLDB |
8 |
5.8755887e-05 |
| 6,608 |
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search |
2014 |
SIGMOD |
14 |
5.73402e-05 |
| 7,032 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
7 |
5.6147056e-05 |
| 7,716 |
Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases |
2013 |
VLDB |
5 |
5.4692136e-05 |
| 7,837 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
7 |
5.4410891e-05 |
| 7,910 |
Near-Duplicate Text Alignment with One Permutation Hashing |
2024 |
SIGMOD |
4 |
5.426303e-05 |
| 8,697 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
5 |
5.2880532e-05 |
| 9,136 |
Balance-Aware Distributed String Similarity-Based Query Processing System |
2019 |
VLDB |
3 |
5.2213971e-05 |
| 9,886 |
META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion |
2016 |
VLDB |
3 |
5.115241e-05 |
| 10,220 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
4 |
5.0572653e-05 |
| 10,489 |
LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH |
2026 |
SIGMOD |
0 |
4.9769913e-05 |
| 10,746 |
Near-Duplicate Text Alignment under Weighted Jaccard Similarity |
2026 |
VLDB |
1 |
4.9769913e-05 |
| 10,897 |
CGIF: Combining Proximity Graphs and Inverted Files for Efficient Filtered Vector Search over Arbitrary Predicates |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 11,856 |
SPINE: Scaling up Programming-by-Negative-Example for String Filtering and Transformation |
2022 |
SIGMOD |
0 |
4.9769913e-05 |