DBScholar

Back to papers

Smurf: Self-Service String Matching Using Random Forests

Summary: Smurf enables self-service string matching with active learning, reducing labeling by 43–76% while maintaining F1. Its RDBMS-style plan optimization reuses computations across RF trees for two string sets, advancing self-service SM and scalable RF over structured data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb70c8269a8d18cba
Venue
VLDB
Year
2019
Pagerank
6.7994518e-05
Overall Rank
4,109 | 72.39%
DOI
10.14778/3291264.3291272
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{c_vldb19,
        title = {{Smurf: Self-Service String Matching Using Random Forests}},
        author = {C., Paul Suganthan G. and Ardalan, Adel and Doan, AnHai and Akella, Aditya},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {3},
        pages = {278--291},
        doi = {10.14778/3291264.3291272},
        url = {https://doi.org/10.14778/3291264.3291272},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 8 of 8 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 25 of 25 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
108 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033040246
129 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.00030395767
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027151132
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025319937
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020001237
433 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018324965
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016923519
530 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00016847532
731 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014400356
827 Adaptive Ordering of Pipelined Stream Filters 2004 SIGMOD 0.00013629035
1,132 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011893781
1,421 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010722146
1,519 SPRINT: A Scalable Parallel Classifier for Data Mining 1996 VLDB 0.00010389098
1,655 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 9.9694129e-05
2,074 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.0821759e-05
2,223 String Similarity Joins: An Experimental Evaluation 2014 VLDB 8.8105347e-05
2,513 An Empirical Evaluation of Set Similarity Join Techniques 2016 VLDB 8.3640659e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2699584e-05
2,605 PLANET: Massively Parallel Learning of Tree Ensembles with MapReduce 2009 VLDB 8.2273571e-05
3,598 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1756405e-05
4,496 Approximate String Joins with Abbreviations 2018 VLDB 6.5741786e-05
5,914 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9472869e-05
7,032 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6147056e-05
9,794 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 5.1248055e-05
12,250 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching 2018 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Semantically Similar Papers