Back to papers
Smurf: Self-Service String Matching Using Random Forests
Summary: Smurf enables self-service string matching with active learning, reducing labeling by 43–76% while maintaining F1. Its RDBMS-style plan optimization reuses computations across RF trees for two string sets, advancing self-service SM and scalable RF over structured data.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 11965
- Venue
- VLDB
- Year
- 2019
- Pagerank
- 6.2142485e-05
- Overall Rank
- 4,397 | 69.45%
- DOI
-
10.14778/3291264.3291272
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 1,914 |
Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks |
2020 |
SIGMOD |
0.00010111859 |
| 4,211 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
6.3495931e-05 |
| 6,552 |
How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses |
2024 |
VLDB |
5.0109216e-05 |
| 6,750 |
Entity Matching Meets Data Science: A Progress Report from the Magellan Project |
2019 |
SIGMOD |
4.936137e-05 |
| 9,362 |
Discovering Top-k Rules using Subjective and Objective Criteria |
2023 |
SIGMOD |
4.3472627e-05 |
| 10,499 |
Incremental Rule Discovery in Response to Parameter Updates |
2025 |
SIGMOD |
4.1905499e-05 |
| 11,090 |
Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching |
2024 |
VLDB |
4.1905499e-05 |
| 11,487 |
Shahin: Faster Algorithms for Generating Explanations for Multiple Predictions |
2021 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 25 of 25 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 125 |
Approximate String Joins in a Database (Almost) for Free |
2001 |
VLDB |
0.00044946098 |
| 179 |
Efficient and Extensible Algorithms for Multi Query Optimization |
2000 |
SIGMOD |
0.00037637319 |
| 248 |
Efficient set joins on similarity predicates |
2004 |
SIGMOD |
0.00030888982 |
| 264 |
Efficient Exact Set-Similarity Joins |
2006 |
VLDB |
0.00029950264 |
| 442 |
Efficient Parallel Set-Similarity Joins Using MapReduce |
2010 |
SIGMOD |
0.00023095823 |
| 641 |
Corleone: Hands-Off Crowdsourcing for Entity Matching |
2014 |
SIGMOD |
0.00018759417 |
| 705 |
Magellan: Toward Building Entity Matching Management Systems |
2016 |
VLDB |
0.00017779048 |
| 832 |
Learning Linear Regression Models over Factorized Joins |
2016 |
SIGMOD |
0.00016089705 |
| 1,041 |
Adaptive Ordering of Pipelined Stream Filters |
2004 |
SIGMOD |
0.00014470785 |
| 1,108 |
SPRINT: A Scalable Parallel Classifier for Data Mining |
1996 |
VLDB |
0.00013931855 |
| 1,172 |
Learning Generalized Linear Models Over Normalized Data |
2015 |
SIGMOD |
0.00013504249 |
| 1,475 |
Efficient Exploitation of Similar Subexpressions for Query Processing |
2007 |
SIGMOD |
0.00011765071 |
| 1,775 |
V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors |
2012 |
VLDB |
0.00010584816 |
| 2,176 |
Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services |
2017 |
SIGMOD |
9.3729351e-05 |
| 2,387 |
Bed-Tree: An All-Purpose Index Structure for String Similarity Search Based on Edit Distance |
2010 |
SIGMOD |
8.9075095e-05 |
| 2,636 |
PLANET: Massively Parallel Learning of Tree Ensembles with MapReduce |
2009 |
VLDB |
8.401513e-05 |
| 2,729 |
String Similarity Joins: An Experimental Evaluation |
2014 |
VLDB |
8.2175463e-05 |
| 3,139 |
ClusterJoin: A Similarity Joins Framework using Map-Reduce |
2014 |
VLDB |
7.4915127e-05 |
| 3,209 |
An Empirical Evaluation of Set Similarity Join Techniques |
2016 |
VLDB |
7.3793885e-05 |
| 4,350 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
6.2576191e-05 |
| 4,682 |
Approximate String Joins with Abbreviations |
2018 |
VLDB |
5.9949001e-05 |
| 6,559 |
Dima: A Distributed In-Memory Similarity-Based Query Processing System |
2017 |
VLDB |
5.00593e-05 |
| 7,106 |
Efficient Similarity Join and Search on Multi-Attribute Data |
2015 |
SIGMOD |
4.8250163e-05 |
| 9,444 |
On-the-Fly Token Similarity Joins in Relational Databases |
2014 |
SIGMOD |
4.3382418e-05 |
| 11,747 |
CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching |
2018 |
VLDB |
4.1905499e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 8,913 |
PromptEM: Prompt-tuning for Low-resource Generalized Entity Matching |
2023 |
VLDB |
4.4229886e-05 |
| 3,469 |
Deep Learning for Blocking in Entity Matching: A Design Space Exploration |
2021 |
VLDB |
7.0629476e-05 |
| 4,033 |
Flexible String Matching Against Large Databases in Practice |
2004 |
VLDB |
6.5125692e-05 |
| 293 |
Deep Learning for Entity Matching: A Design Space Exploration |
2018 |
SIGMOD |
0.00028661817 |
| 11,090 |
Dealing with Acronyms, Abbreviations, and Typos in Real-World Entity Matching |
2024 |
VLDB |
4.1905499e-05 |
| 11,253 |
Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests |
2023 |
VLDB |
4.1905499e-05 |
| 2,176 |
Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services |
2017 |
SIGMOD |
9.3729351e-05 |
| 2,758 |
A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching |
2020 |
SIGMOD |
8.1668285e-05 |
| 5,872 |
Demonstration of Panda: A Weakly Supervised Entity Matching System |
2021 |
VLDB |
5.2908178e-05 |
| 9,415 |
Ground Truth Inference for Weakly Supervised Entity Matching |
2023 |
SIGMOD |
4.3399748e-05 |