| 353 |
Efficient Parallel Set-Similarity Joins Using MapReduce |
2010 |
SIGMOD |
0.00020497819 |
| 538 |
On Active Learning of Record Matching Packages |
2010 |
SIGMOD |
0.00016945549 |
| 557 |
Management of Probabilistic Data: Foundations and Challenges |
2007 |
PODS |
0.00016599507 |
| 771 |
JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes |
2019 |
SIGMOD |
0.00014215821 |
| 850 |
Framework for Evaluating Clustering Algorithms in Duplicate Detection |
2009 |
VLDB |
0.00013661662 |
| 887 |
Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints |
2008 |
VLDB |
0.00013452998 |
| 1,012 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012747701 |
| 1,279 |
Entity Matching: How Similar Is Similar |
2011 |
VLDB |
0.00011448343 |
| 1,403 |
V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors |
2012 |
VLDB |
0.00010985239 |
| 1,548 |
Example-driven Design of Efficient Record Matching Queries |
2007 |
VLDB |
0.0001044945 |
| 1,676 |
Extending Autocompletion To Tolerate Errors |
2009 |
SIGMOD |
0.00010121328 |
| 1,938 |
ATLAS: A Probabilistic Algorithm for High Dimensional Similarity Search |
2011 |
SIGMOD |
9.5248884e-05 |
| 2,008 |
Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance |
2010 |
SIGMOD |
9.3978374e-05 |
| 2,037 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
9.3510621e-05 |
| 2,158 |
String Similarity Joins: An Experimental Evaluation |
2014 |
VLDB |
9.1187342e-05 |
| 2,219 |
WHAM: A High-throughput Sequence Alignment Method |
2011 |
SIGMOD |
8.9998232e-05 |
| 2,291 |
Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries |
2008 |
VLDB |
8.8798097e-05 |
| 2,433 |
ZeroER: Entity Resolution using Zero Labeled Examples |
2020 |
SIGMOD |
8.6525126e-05 |
| 2,534 |
ClusterJoin: A Similarity Joins Framework using Map-Reduce |
2014 |
VLDB |
8.520571e-05 |
| 2,683 |
An Empirical Evaluation of Set Similarity Join Techniques |
2016 |
VLDB |
8.3233182e-05 |
| 2,969 |
Spatio-Textual Similarity Joins |
2013 |
VLDB |
7.9604274e-05 |
| 2,971 |
Leveraging Set Relations in Exact Set Similarity Join |
2017 |
VLDB |
7.95959e-05 |
| 3,156 |
Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme |
2011 |
SIGMOD |
7.7577179e-05 |
| 3,329 |
Benchmarking Declarative Approximate Selection Predicates |
2007 |
SIGMOD |
7.5731286e-05 |
| 3,397 |
An Efficient Filter for Approximate Membership Checking |
2008 |
SIGMOD |
7.5171701e-05 |
| 3,399 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
7.5167701e-05 |
| 3,401 |
Efficient Approximate Entity Extraction with Edit Distance Constraints |
2009 |
SIGMOD |
7.5137307e-05 |
| 3,697 |
Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints |
2010 |
VLDB |
7.2570567e-05 |
| 3,702 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
7.2540345e-05 |
| 3,966 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
7.0599892e-05 |
| 3,995 |
Smurf: Self-Service String Matching Using Random Forests |
2019 |
VLDB |
7.0412373e-05 |
| 4,204 |
Set Similarity Joins on MapReduce: An Experimental Survey |
2018 |
VLDB |
6.9001164e-05 |
| 4,312 |
Incremental Maintenance of Length Normalized Indexes for Approximate String Matching |
2009 |
SIGMOD |
6.8382796e-05 |
| 4,333 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
6.8301851e-05 |
| 4,418 |
Power-Law Based Estimation of Set Similarity Join Size |
2009 |
VLDB |
6.7764944e-05 |
| 4,618 |
On the Complexity of Inner Product Similarity Join |
2016 |
PODS |
6.6665536e-05 |
| 4,641 |
SilkMoth: An Efficient Method for Finding Related Sets with Maximum Matching Constraints |
2017 |
VLDB |
6.6531306e-05 |
| 4,821 |
Similarity Join Size Estimation using Locality Sensitive Hashing |
2011 |
VLDB |
6.5602214e-05 |
| 4,919 |
SEAL: Spatio-Textual Similarity Search |
2012 |
VLDB |
6.5100672e-05 |
| 4,943 |
Scalable Ad-hoc Entity Extraction from Text Collections |
2008 |
VLDB |
6.4989078e-05 |
| 5,072 |
String Similarity Measures and Joins with Synonyms |
2013 |
SIGMOD |
6.4430928e-05 |
| 5,088 |
Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples |
2021 |
SIGMOD |
6.4348674e-05 |
| 5,210 |
On Link-based Similarity Join |
2011 |
VLDB |
6.3856427e-05 |
| 5,263 |
On Indexing Error-Tolerant Set Containment |
2010 |
SIGMOD |
6.3625471e-05 |
| 5,347 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
6.3242768e-05 |
| 5,595 |
Efficient Approximate Search on String Collections (Tutorial) |
2009 |
VLDB |
6.2222634e-05 |
| 5,597 |
Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition |
2018 |
VLDB |
6.221166e-05 |
| 5,605 |
Pigeonring: A Principle for Faster Thresholded Similarity Search |
2019 |
VLDB |
6.2161107e-05 |
| 5,898 |
The Communication Complexity of Distributed Set-Joins with Applications to Matrix Multiplication |
2015 |
PODS |
6.1115843e-05 |
| 6,202 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
6.0164826e-05 |