Cleaning Crowdsourced Labels Using Oracles for Statistical Classification
Summary: Oracle-based label cleaning for crowdsourced data in classification. TARS estimates test performance from noisy labels with confidence bounds and selects which labels to clean to boost training accuracy under budget, beating existing strategies. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Mohamad Dolatshah (Simon Fraser University)
- 2. Mathew Teoh (Simon Fraser University)
- 3. Jiannan Wang (Simon Fraser University)
- 4. Jian Pei (JD Group; Simon Fraser University)
BibTeX Citation
@article{dolatshah_vldb19,
title = {{Cleaning Crowdsourced Labels Using Oracles for Statistical Classification}},
author = {Dolatshah, Mohamad and Teoh, Mathew and Wang, Jiannan and Pei, Jian},
journal = {PVLDB},
series = {{VLDB} '19},
volume = {12},
number = {4},
pages = {376--389},
doi = {10.14778/3297753.3297758},
url = {https://doi.org/10.14778/3297753.3297758},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 22 of 22 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,785 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 2 | 7,665 | Towards Globally Optimal Crowdsourcing Quality Management: The Uniform Worker Setting | 2016 | SIGMOD |
| 3 | 10,604 | Fault Lines: Benchmarking the Impact of Label Data Quality on ML Robustness and Fairness | 2026 | VLDB |
| 4 | 2,570 | Query-Oriented Data Cleaning with Oracles | 2015 | SIGMOD |
| 5 | 11,343 | Generalizable Data Cleaning of Tabular Data in Latent Space | 2024 | VLDB |
| 6 | 6,908 | Cost-Effective Data Annotation using Game-Based Crowdsourcing | 2019 | VLDB |
| 7 | 4,821 | Crowdsourced Top-k Queries by Confidence-Aware Pairwise Judgments | 2017 | SIGMOD |
| 8 | 11,143 | k-Clustering with Comparison and Distance Oracles | 2024 | PODS |
| 9 | 2,626 | Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning | 2015 | VLDB |
| 10 | 9,820 | How to Design Robust Algorithms using Noisy Comparison Oracle | 2021 | VLDB |