DBScholar

Back to papers

Cleaning Crowdsourced Labels Using Oracles for Statistical Classification

Summary: Oracle-based label cleaning for crowdsourced data in classification. TARS estimates test performance from noisy labels with confidence bounds and selects which labels to clean to boost training accuracy under budget, beating existing strategies. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12160
Venue
VLDB
Year
2019
Pagerank
7.5706653e-05
Overall Rank
3,281 | 77.50%
DOI
10.14778/3297753.3297758

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{dolatshah_vldb19,
        title = {{Cleaning Crowdsourced Labels Using Oracles for Statistical Classification}},
        author = {Dolatshah, Mohamad and Teoh, Mathew and Wang, Jiannan and Pei, Jian},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {4},
        pages = {376--389},
        doi = {10.14778/3297753.3297758},
        url = {https://doi.org/10.14778/3297753.3297758},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 12 of 12 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 22 of 22 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
196 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025780596
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
714 Guided Data Repair 2011 VLDB 0.00014662041
933 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00013111293
1,101 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012168934
1,270 CDAS: A Crowdsourcing Data Analytics System 2012 VLDB 0.00011388453
1,323 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00011152602
1,643 Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services 2017 SIGMOD 0.00010134956
1,736 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.8984415e-05
1,783 Data Fusion – Resolving Data Conflicts for Integration 2009 VLDB 9.772895e-05
2,570 Query-Oriented Data Cleaning with Oracles 2015 SIGMOD 8.4061164e-05
2,626 Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning 2015 VLDB 8.3291889e-05
2,636 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.31868e-05
3,231 Truth Inference in Crowdsourcing: Is the Problem Solved? 2017 VLDB 7.6191767e-05
3,262 iCrowd: An Adaptive Crowdsourcing Framework 2015 SIGMOD 7.5846052e-05
3,602 Online Entity Resolution Using an Oracle 2016 VLDB 7.2691521e-05
3,715 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.1763559e-05
3,904 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 7.0304212e-05
3,971 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.9835263e-05
5,143 An Online Cost Sensitive Decision-Making Method in Crowdsourcing Systems 2013 SIGMOD 6.3476753e-05
5,981 Truth Discovery and Crowdsourcing Aggregation: A Unified Perspective 2015 VLDB 6.0211322e-05
8,351 Minimizing Efforts in Validating Crowd Answers 2015 SIGMOD 5.4454221e-05
Previous Page 1 / 1 Next

Semantically Similar Papers