DBScholar

Back to papers

Smurf: Self-Service String Matching Using Random Forests

Summary: Smurf enables self-service string matching with active learning, reducing labeling by 43–76% while maintaining F1. Its RDBMS-style plan optimization reuses computations across RF trees for two string sets, advancing self-service SM and scalable RF over structured data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12152
Venue
VLDB
Year
2019
Pagerank
6.949387e-05
Overall Rank
4,023 | 72.40%
DOI
10.14778/3291264.3291272

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{c_vldb19,
        title = {{Smurf: Self-Service String Matching Using Random Forests}},
        author = {C., Paul Suganthan G. and Ardalan, Adel and Doan, AnHai and Akella, Aditya},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {3},
        pages = {278--291},
        doi = {10.14778/3291264.3291272},
        url = {https://doi.org/10.14778/3291264.3291272},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 8 of 8 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 25 of 25 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
107 Approximate String Joins in a Database (Almost) for Free 2001 VLDB 0.00033511706
128 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.0003072825
169 Efficient Exact Set-Similarity Joins 2006 VLDB 0.0002743469
200 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025597287
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
439 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018464913
529 Magellan: Toward Building Entity Matching Management Systems 2016 VLDB 0.00017096361
536 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.0001693369
715 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014655327
813 Adaptive Ordering of Pipelined Stream Filters 2004 SIGMOD 0.00013846487
1,154 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011934202
1,415 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010840141
1,485 SPRINT: A Scalable Parallel Classifier for Data Mining 1996 VLDB 0.00010628998
1,643 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 0.00010134956
2,036 Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance 2010 SIGMOD 9.2782094e-05
2,186 String Similarity Joins: An Experimental Evaluation 2014 VLDB 9.0001436e-05
2,501 An Empirical Evaluation of Set Similarity Join Techniques 2016 VLDB 8.4975661e-05
2,560 PLANET: Massively Parallel Learning of Tree Ensembles with MapReduce 2009 VLDB 8.4143663e-05
2,567 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.4098241e-05
3,724 Overlap Set Similarity Joins with Theoretical Guarantees 2018 SIGMOD 7.1715735e-05
4,396 Approximate String Joins with Abbreviations 2018 VLDB 6.7268636e-05
6,290 Dima: A Distributed In-Memory Similarity-Based Query Processing System 2017 VLDB 5.9253163e-05
6,906 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.7418509e-05
9,613 On-the-Fly Token Similarity Joins in Relational Databases 2014 SIGMOD 5.2447096e-05
11,945 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching 2018 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers