| 779 |
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes |
2019 |
SIGMOD |
0.00014092047 |
| 963 |
The Data Civilizer System |
2017 |
CIDR |
0.00012935145 |
| 1,351 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00011064851 |
| 1,886 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
9.5358137e-05 |
| 2,303 |
SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search |
2024 |
SIGMOD |
8.7783079e-05 |
| 2,572 |
DeltaPQ: Lossless Product Quantization Code Compression for High Dimensional Similarity Search |
2020 |
VLDB |
8.4027322e-05 |
| 3,474 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
7.3859271e-05 |
| 3,724 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
7.1715735e-05 |
| 4,396 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
6.7268636e-05 |
| 4,538 |
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach |
2016 |
SIGMOD |
6.6404776e-05 |
| 4,707 |
SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
6.552423e-05 |
| 5,317 |
Efficient Load-Balanced Butterfly Counting on GPU |
2022 |
VLDB |
6.2685847e-05 |
| 5,412 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
6.2272563e-05 |
| 5,644 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
6.1378898e-05 |
| 5,878 |
Distributed Graph Simulation: Impossibility and Possibility |
2014 |
VLDB |
6.0539311e-05 |
| 6,290 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
5.9253163e-05 |
| 6,404 |
ARKGraph: All-Range Approximate K-Nearest-Neighbor Graph |
2023 |
VLDB |
5.8859287e-05 |
| 6,484 |
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search |
2014 |
SIGMOD |
5.8665833e-05 |
| 6,906 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
5.7418509e-05 |
| 7,426 |
Dynamic Range-Filtering Approximate Nearest Neighbor Search |
2025 |
VLDB |
5.6211019e-05 |
| 7,575 |
Scalable Column Concept Determination for Web Tables Using Large Knowledge Bases |
2013 |
VLDB |
5.5937684e-05 |
| 7,676 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
5.5686107e-05 |
| 7,745 |
Near-Duplicate Text Alignment with One Permutation Hashing |
2024 |
SIGMOD |
5.5534781e-05 |
| 8,522 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
5.4119882e-05 |
| 9,705 |
META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion |
2016 |
VLDB |
5.2351259e-05 |
| 9,979 |
Balance-Aware Distributed String Similarity-Based Query Processing System |
2019 |
VLDB |
5.1845938e-05 |
| 10,024 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
5.1757914e-05 |
| 10,265 |
LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH |
2026 |
SIGMOD |
5.093636e-05 |
| 10,554 |
Near-Duplicate Text Alignment under Weighted Jaccard Similarity |
2026 |
VLDB |
5.093636e-05 |
| 11,541 |
SPINE: Scaling up Programming-by-Negative-Example for String Filtering and Transformation |
2022 |
SIGMOD |
5.093636e-05 |