DBScholar

Back to papers

Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services

Summary: Falcon scales hands-off crowdsourced EM beyond Corleone with RDBMS-style planning on Hadoop. It defines EM operators, turns workflows into executable plans mixing machine and crowd tasks, using crowd time to mask machine time for million-tuple cloud-scale EM. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc07e9b382b12317b
Venue
SIGMOD
Year
2017
Pagerank
9.9739611e-05
Overall Rank
1,654 | 88.89%
DOI
10.1145/3035918.3035960

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{das_sigmod17,
        title = {{Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services}},
        author = {Das, Sanjib and C., Paul Suganthan G. and Doan, AnHai and Naughton, Jeffrey F. and Krishnan, Ganesh and Deep, Rohit and Arcaute, Esteban and Raghavendra, Vijay and Park, Youngchoon},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3035960},
        url = {https://dl.acm.org/doi/10.1145/3035918.3035960},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 30 of 30 citing papers.

Rank Citing Paper Year Venue Pagerank
457 Distributed Representations of Tuples for Entity Resolution 2018 VLDB 0.00017907103
1,391 Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks 2020 SIGMOD 0.00010816237
2,475 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.410678e-05
2,514 Deep Learning for Blocking in Entity Matching: A Design Space Exploration 2021 VLDB 8.3648432e-05
3,230 Similarity Query Processing for High-Dimensional Data 2020 VLDB 7.5056037e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
3,473 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.2728706e-05
4,107 Smurf: Self-Service String Matching Using Random Forests 2019 VLDB 6.8026037e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
5,049 BEER: Blocking for Effective Entity Resolution 2021 SIGMOD 6.297901e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,126 Monotonic Cardinality Estimation of Similarity Selection: A Deep Learning Approach 2020 SIGMOD 6.261175e-05
6,300 Sparkly: A Simple yet Surprisingly Strong TF/IDF Blocker for Entity Matching 2023 VLDB 5.8189161e-05
6,397 Human-in-the-loop Data Integration 2017 VLDB 5.7989499e-05
6,525 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.756805e-05
7,052 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.6127577e-05
7,415 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5332662e-05
7,509 Entity Matching Meets Data Science: A Progress Report from the Magellan Project 2019 SIGMOD 5.5062607e-05
7,872 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4362062e-05
8,276 Consistent and Flexible Selectivity Estimation for High-Dimensional Data 2021 SIGMOD 5.3641556e-05
8,282 Online Topic-Aware Entity Resolution Over Incomplete Data Streams 2021 SIGMOD 5.3629192e-05
9,126 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.22387e-05
9,241 Deep Active Alignment of Knowledge Graph Entities and Schemata 2023 SIGMOD 5.2056825e-05
9,458 HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs 2025 VLDB 5.173446e-05
9,565 Deduplicated Sampling On-Demand 2025 VLDB 5.1571823e-05
9,806 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1257999e-05
10,527 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 4.9793485e-05
11,738 Splitting Tuples of Mismatched Entities 2023 SIGMOD 4.9793485e-05
11,744 VersaMatch: Ontology Matching with Weak Supervision 2023 VLDB 4.9793485e-05
12,244 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching 2018 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 30 of 30 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
92 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034672523
198 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025555196
201 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025331535
204 Declarative Data Cleaning: Language, Model, and Algorithms 2001 VLDB 0.00025190386
257 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00022962347
259 Answering Queries using Humans, Algorithms and Databases 2011 CIDR 0.00022923243
266 Human-powered Sorts and Joins 2012 VLDB 0.00022739124
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
433 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018332741
749 So Who Won? Dynamic Max Discovery with the Crowd 2012 SIGMOD 0.00014265279
787 Human-Assisted Graph Search: It’s Okay to Ask Questions 2011 VLDB 0.00013988274
865 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.0001338765
871 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013343705
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013020115
933 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00013004422
987 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012660627
1,472 Crowdsourcing Algorithms for Entity Resolution 2014 VLDB 0.00010558263
1,926 Counting with the Crowd 2013 VLDB 9.3637786e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2738285e-05
2,642 Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning 2015 VLDB 8.1778168e-05
2,832 Crowd Mining 2013 SIGMOD 7.9590386e-05
2,948 Distributed Data Deduplication 2016 VLDB 7.8230494e-05
3,350 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3910669e-05
3,816 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.0072547e-05
4,024 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.8472018e-05
4,623 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.4978811e-05
6,668 Query Optimization over Crowdsourced Data 2013 VLDB 5.7147473e-05
7,030 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.6173605e-05
8,777 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.2792451e-05
Previous Page 1 / 1 Next

Semantically Similar Papers