DBScholar

Back to papers

Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services

Summary: Falcon scales hands-off crowdsourced EM beyond Corleone with RDBMS-style planning on Hadoop. It defines EM operators, turns workflows into executable plans mixing machine and crowd tasks, using crowd time to mask machine time for million-tuple cloud-scale EM. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5388
Venue
SIGMOD
Year
2017
Pagerank
0.00010134956
Overall Rank
1,643 | 88.73%
DOI
10.1145/3035918.3035960

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{das_sigmod17,
        title = {{Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services}},
        author = {Das, Sanjib and C., Paul Suganthan G. and Doan, AnHai and Naughton, Jeffrey F. and Krishnan, Ganesh and Deep, Rohit and Arcaute, Esteban and Raghavendra, Vijay and Park, Youngchoon},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3035960},
        url = {https://dl.acm.org/doi/10.1145/3035918.3035960},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 30 of 30 citing papers.

Rank Citing Paper Year Venue Pagerank
489 Distributed Representations of Tuples for Entity Resolution 2018 VLDB 0.0001761456
1,402 Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks 2020 SIGMOD 0.00010888094
2,463 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.5486912e-05
2,475 Deep Learning for Blocking in Entity Matching: A Design Space Exploration 2021 VLDB 8.5277654e-05
3,281 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.5706653e-05
3,310 Similarity Query Processing for High-Dimensional Data 2020 VLDB 7.5363562e-05
3,436 Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration 2023 SIGMOD 7.4157897e-05
4,023 Smurf: Self-Service String Matching Using Random Forests 2019 VLDB 6.949387e-05
4,100 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.9016092e-05
4,940 BEER: Blocking for Effective Entity Resolution 2021 SIGMOD 6.4336209e-05
4,966 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.4225454e-05
5,011 Monotonic Cardinality Estimation of Similarity Selection: A Deep Learning Approach 2020 SIGMOD 6.4020848e-05
6,167 Sparkly: A Simple yet Surprisingly Strong TF/IDF Blocker for Entity Matching 2023 VLDB 5.9524736e-05
6,403 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.8859374e-05
6,908 Cost-Effective Data Annotation using Game-Based Crowdsourcing 2019 VLDB 5.7415834e-05
7,373 Entity Matching Meets Data Science: A Progress Report from the Magellan Project 2019 SIGMOD 5.6311132e-05
7,424 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.6214566e-05
7,493 Human-in-the-loop Data Integration 2017 VLDB 5.6046905e-05
8,105 Online Topic-Aware Entity Resolution Over Incomplete Data Streams 2021 SIGMOD 5.4860105e-05
8,124 Consistent and Flexible Selectivity Estimation for High-Dimensional Data 2021 SIGMOD 5.4829513e-05
9,063 Deep Active Alignment of Knowledge Graph Entities and Schemata 2023 SIGMOD 5.3251649e-05
9,381 Deduplicated Sampling On-Demand 2025 VLDB 5.2755515e-05
9,627 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.2434488e-05
9,647 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.2430158e-05
9,979 Balance-Aware Distributed String Similarity-Based Query Processing System 2019 VLDB 5.1845938e-05
9,997 HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUs 2025 VLDB 5.1814573e-05
10,318 In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration 2026 SIGMOD 5.093636e-05
11,424 Splitting Tuples of Mismatched Entities 2023 SIGMOD 5.093636e-05
11,430 VersaMatch: Ontology Matching with Weak Supervision 2023 VLDB 5.093636e-05
11,945 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching 2018 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 30 of 30 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
90 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034951786
196 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025780596
200 Efficient set joins on similarity predicates 2004 SIGMOD 0.00025597287
201 Declarative Data Cleaning: Language, Model, and Algorithms 2001 VLDB 0.00025558602
250 Answering Queries using Humans, Algorithms and Databases 2011 CIDR 0.00023261164
251 Crowdsourced Databases: Query Processing with People 2011 CIDR 0.00023261113
265 Human-powered Sorts and Joins 2012 VLDB 0.00022935368
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
439 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018464913
743 So Who Won? Dynamic Max Discovery with the Crowd 2012 SIGMOD 0.00014421358
767 Human-Assisted Graph Search: It’s Okay to Ask Questions 2011 VLDB 0.00014208622
843 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.00013666161
852 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013604253
911 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013283031
933 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00013111293
997 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012755983
1,443 Crowdsourcing Algorithms for Entity Resolution 2014 VLDB 0.00010773106
1,910 Counting with the Crowd 2013 VLDB 9.4972788e-05
2,398 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.631172e-05
2,567 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.4098241e-05
2,626 Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning 2015 VLDB 8.3291889e-05
2,789 Crowd Mining 2013 SIGMOD 8.1203892e-05
2,893 Distributed Data Deduplication 2016 VLDB 7.983961e-05
3,474 An Efficient Partition Based Method for Exact Set Similarity Joins 2016 VLDB 7.3859271e-05
3,731 Trie-Join: Efficient Trie-based String Similarity Joins with Edit-Distance Constraints 2010 VLDB 7.1663952e-05
3,971 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.9835263e-05
4,538 Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach 2016 SIGMOD 6.6404776e-05
6,548 Query Optimization over Crowdsourced Data 2013 VLDB 5.8443704e-05
6,906 Efficient Similarity Join and Search on Multi-Attribute Data 2015 SIGMOD 5.7418509e-05
8,620 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.3991092e-05
Previous Page 1 / 1 Next

Semantically Similar Papers