DBScholar

Back to papers

Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning

Summary: Scalable active-learning for crowd-sourced databases, combining ML with human labeling via nonparametric bootstrap. MTurk and 15 datasets show 1–2 orders of magnitude fewer questions than baselines and 4.5–44× faster than prior AL. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hf18fc83ae7e33252
Venue
VLDB
Year
2015
Pagerank
8.1778168e-05
Overall Rank
2,642 | 82.24%
DOI
10.14778/2735471.2735474

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{mozafari_vldb15,
        title = {{Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning}},
        author = {Mozafari, Barzan and Sarkar, Purna and Franklin, Michael and Jordan, Michael and Madden, Samuel},
        journal = {PVLDB},
        series = {{VLDB} '15},
        volume = {8},
        number = {2},
        pages = {125},
        doi = {10.14778/2735471.2735474},
        url = {https://doi.org/10.14778/2735471.2735474},
        year = {2015}
}

Incoming Citations (Sorted by Pagerank)

Showing 18 of 18 citing papers.

Rank Citing Paper Year Venue Pagerank
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
784 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014012614
1,043 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012335114
1,654 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 9.9739611e-05
2,275 Active Learning for ML Enhanced Database Systems 2020 SIGMOD 8.7090584e-05
2,475 A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching 2020 SIGMOD 8.410678e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
4,024 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.8472018e-05
4,939 Deep Indexed Active Learning for Matching Heterogeneous Entity Representations 2022 VLDB 6.3458427e-05
5,915 In Search of an Entity Resolution OASIS: Optimal Asymptotic Sequential Importance Sampling 2017 VLDB 5.948779e-05
7,500 Crowdsourced Data Management: Overview and Challenges 2017 SIGMOD 5.5083793e-05
7,778 User Guidance for Efficient Fact Checking 2019 VLDB 5.4546499e-05
9,785 The Battleship Approach to the Low Resource Entity Matching Problem 2023 SIGMOD 5.1283279e-05
10,244 Towards Interpretable and Learnable Risk Analysis for Entity Resolution 2020 SIGMOD 5.0525742e-05
10,748 ALER: An Active Learning Hybrid System for Efficient Entity Resolution 2026 VLDB 4.9793485e-05
11,744 VersaMatch: Ontology Matching with Weak Supervision 2023 VLDB 4.9793485e-05
12,090 Recommending Deployment Strategies for Collaborative Tasks 2020 SIGMOD 4.9793485e-05
12,273 Staging User Feedback toward Rapid Conflict Resolution in Data Fusion 2017 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
92 CrowdDB: Answering Queries with Crowdsourcing 2011 SIGMOD 0.00034672523
198 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025555196
266 Human-powered Sorts and Joins 2012 VLDB 0.00022739124
534 On Active Learning of Record Matching Packages 2010 SIGMOD 0.0001680637
987 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012660627
1,916 The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing 2014 SIGMOD 9.3837729e-05
1,926 Counting with the Crowd 2013 VLDB 9.3637786e-05
5,473 ABS: a System for Scalable Approximate Queries with Accuracy Guarantees 2014 SIGMOD 6.1178467e-05
Previous Page 1 / 1 Next

Semantically Similar Papers