| 360 |
Efficient Parallel Set-Similarity Joins Using MapReduce |
2010 |
SIGMOD |
0.00020009936 |
| 534 |
On Active Learning of Record Matching Packages |
2010 |
SIGMOD |
0.0001680637 |
| 587 |
Management of Probabilistic Data: Foundations and Challenges |
2007 |
PODS |
0.00015933201 |
| 694 |
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes |
2019 |
SIGMOD |
0.00014727089 |
| 885 |
Framework for Evaluating Clustering Algorithms in Duplicate Detection |
2009 |
VLDB |
0.00013260551 |
| 929 |
Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints |
2008 |
VLDB |
0.00013020115 |
| 963 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012816649 |
| 1,247 |
Entity Matching: How Similar Is Similar |
2011 |
VLDB |
0.00011346515 |
| 1,421 |
V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors |
2012 |
VLDB |
0.00010726757 |
| 1,581 |
Example-driven Design of Efficient Record Matching Queries |
2007 |
VLDB |
0.00010180038 |
| 1,737 |
Extending Autocompletion To Tolerate Errors |
2009 |
SIGMOD |
9.7494976e-05 |
| 1,923 |
ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search |
2011 |
SIGMOD |
9.3750526e-05 |
| 1,933 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
9.3459285e-05 |
| 2,072 |
Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance |
2010 |
SIGMOD |
9.0864612e-05 |
| 2,220 |
String Similarity Joins: An Experimental Evaluation |
2014 |
VLDB |
8.8146984e-05 |
| 2,323 |
ZeroER: Entity Resolution using Zero Labeled Examples |
2020 |
SIGMOD |
8.6348884e-05 |
| 2,354 |
Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries |
2008 |
VLDB |
8.5896515e-05 |
| 2,508 |
WHAM: A High-throughput Sequence Alignment Method |
2011 |
SIGMOD |
8.371338e-05 |
| 2,513 |
An Empirical Evaluation of Set Similarity Join Techniques |
2016 |
VLDB |
8.3679178e-05 |
| 2,577 |
ClusterJoin: A Similarity Joins Framework using Map-Reduce |
2014 |
VLDB |
8.2738285e-05 |
| 2,921 |
Leveraging Set Relations in Exact Set Similarity Join |
2017 |
VLDB |
7.8519256e-05 |
| 2,944 |
Spatio-Textual Similarity Joins |
2013 |
VLDB |
7.828107e-05 |
| 3,255 |
Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme |
2011 |
SIGMOD |
7.4890984e-05 |
| 3,350 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
7.3910669e-05 |
| 3,403 |
Benchmarking Declarative Approximate Selection Predicates |
2007 |
SIGMOD |
7.3300953e-05 |
| 3,501 |
An Efficient Filter for Approximate Membership Checking |
2008 |
SIGMOD |
7.2531528e-05 |
| 3,512 |
Efficient Approximate Entity Extraction with Edit Distance Constraints |
2009 |
SIGMOD |
7.2439455e-05 |
| 3,596 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
7.1790375e-05 |
| 3,809 |
On the Complexity of Inner Product Similarity Join |
2016 |
PODS |
7.0089846e-05 |
| 3,816 |
Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints |
2010 |
VLDB |
7.0072547e-05 |
| 4,107 |
Smurf: Self-Service String Matching Using Random Forests |
2019 |
VLDB |
6.8026037e-05 |
| 4,147 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
6.7803631e-05 |
| 4,274 |
Set Similarity Joins on MapReduce: An Experimental Survey |
2018 |
VLDB |
6.6918599e-05 |
| 4,439 |
Incremental Maintenance of Length Normalized Indexes for Approximate String Matching |
2009 |
SIGMOD |
6.5998035e-05 |
| 4,493 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
6.5772891e-05 |
| 4,536 |
Power-Law Based Estimation of Set Similarity Join Size |
2009 |
VLDB |
6.5557137e-05 |
| 4,571 |
SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
6.5286445e-05 |
| 4,736 |
Scalable Ad-hoc Entity Extraction from Text Collections |
2008 |
VLDB |
6.4451961e-05 |
| 4,958 |
Similarity Join Size Estimation using Locality Sensitive Hashing |
2011 |
VLDB |
6.3394776e-05 |
| 4,970 |
Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples |
2021 |
SIGMOD |
6.3329595e-05 |
| 5,112 |
SEAL: Spatio-Textual Similarity Search |
2012 |
VLDB |
6.268442e-05 |
| 5,207 |
String Similarity Measures and Joins with Synonyms |
2013 |
SIGMOD |
6.2272365e-05 |
| 5,309 |
On Link-based Similarity Join |
2011 |
VLDB |
6.1865348e-05 |
| 5,333 |
On Indexing Error-Tolerant Set Containment |
2010 |
SIGMOD |
6.1750986e-05 |
| 5,444 |
Pigeonring: A Principle for Faster Thresholded Similarity Search |
2019 |
VLDB |
6.1268526e-05 |
| 5,549 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
6.088104e-05 |
| 5,764 |
Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition |
2018 |
VLDB |
6.0030951e-05 |
| 5,795 |
Efficient Approximate Search on String Collections (Tutorial) |
2009 |
VLDB |
5.9926463e-05 |
| 5,912 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
5.9501002e-05 |
| 6,121 |
The Communication Complexity of Distributed Set-Joins with Applications to Matrix Multiplication |
2015 |
PODS |
5.8800995e-05 |