DBScholar

Back to papers

Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services

Summary: Falcon scales hands-off crowdsourced EM beyond Corleone with RDBMS-style planning on Hadoop. It defines EM operators, turns workflows into executable plans mixing machine and crowd tasks, using crowd time to mask machine time for million-tuple cloud-scale EM. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc07e9b382b12317b
Venue
SIGMOD
Year
2017
Pagerank
9.9694129e-05
Overall Rank
1,655 | 88.88%
DOI
10.1145/3035918.3035960

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{das_sigmod17,
        title = {{Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services}},
        author = {Das, Sanjib and C., Paul Suganthan G. and Doan, AnHai and Naughton, Jeffrey F. and Krishnan, Ganesh and Deep, Rohit and Arcaute, Esteban and Raghavendra, Vijay and Park, Youngchoon},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3035960},
        url = {https://dl.acm.org/doi/10.1145/3035918.3035960},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 30 of 30 citing papers.

Rank Citing Paper Year Venue Pagerank
457 Distributed Representations of Tuples for Entity Resolution 2018 VLDB 0.00017899824
1,391 Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks 2020 SIGMOD 0.00010812249
2,475 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.4074448e-05
2,514 Deep Learning for Blocking in Entity Matching: A Design Space Exploration 2021 VLDB 8.3610263e-05
3,232 Similarity Query Processing for High-Dimensional Data 2020 VLDB 7.5029541e-05
3,308 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4379101e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2697311e-05
4,109 Smurf: Self-Service String Matching Using Random Forests 2019 VLDB 6.7994518e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7576159e-05
5,052 BEER: Blocking for Effective Entity Resolution 2021 SIGMOD 6.2949299e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,129 Monotonic Cardinality Estimation of Similarity Selection: A Deep Learning Approach 2020 SIGMOD 6.2582129e-05
6,304 Sparkly: A Simple yet Surprisingly Strong TF/IDF Blocker for Entity Matching 2023 VLDB 5.8161615e-05
6,400 Human-in-the-loop Data Integration 2017 VLDB 5.7962311e-05
6,528 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.7540798e-05
7,054 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.6101007e-05
7,418 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5306468e-05
7,514 Entity Matching Meets Data Science: A Progress Report from the Magellan Project 2019 SIGMOD 5.5036541e-05
7,876 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4336606e-05
8,282 Consistent and Flexible Selectivity Estimation for High-Dimensional Data 2021 SIGMOD 5.3616162e-05
8,288 Online Topic-Aware Entity Resolution Over Incomplete Data Streams 2021 SIGMOD 5.3603804e-05
9,136 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.2213971e-05
9,251 Deep Active Alignment of Knowledge Graph Entities and Schemata 2023 SIGMOD 5.2032182e-05
9,467 HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs 2025 VLDB 5.1709969e-05
9,573 Deduplicated Sampling On-Demand 2025 VLDB 5.154741e-05
9,813 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1233734e-05
10,538 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 4.9769913e-05
11,744 Splitting Tuples of Mismatched Entities 2023 SIGMOD 4.9769913e-05
11,750 VersaMatch: Ontology Matching with Weak Supervision 2023 VLDB 4.9769913e-05
12,250 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching 2018 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 30 of 30 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
92 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034670735
198 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025546182
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025319937
204 Declarative Data Cleaning: Language, Model, and Algorithms 2001 VLDB 0.00025179068
257 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00022958404
259 Answering Queries using Humans, Algorithms and Databases 2011 CIDR 0.00022916014
266 Human-powered Sorts and Joins 2012 VLDB 0.00022735106
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020001237
433 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018324965
750 So Who Won? Dynamic Max Discovery with the Crowd 2012 SIGMOD 0.00014258988
787 Human-Assisted Graph Search: It’s Okay to Ask Questions 2011 VLDB 0.00013982951
866 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.00013381756
871 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013338722
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013014029
934 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00012998402
987 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012655094
1,472 Crowdsourcing Algorithms for Entity Resolution 2014 VLDB 0.00010553304
1,926 Counting with the Crowd 2013 VLDB 9.3636304e-05
2,424 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.483813e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2699584e-05
2,643 Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning 2015 VLDB 8.1741373e-05
2,832 Crowd Mining 2013 SIGMOD 7.9554419e-05
2,949 Distributed Data Deduplication 2016 VLDB 7.8193962e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3875743e-05
3,817 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0039415e-05
4,025 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.8439934e-05
4,625 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.4948389e-05
6,672 Query Optimization over Crowdsourced Data 2013 VLDB 5.7120461e-05
7,032 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6147056e-05
8,785 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.2767488e-05
Previous Page 1 / 1 Next

Semantically Similar Papers