Back to papers
MisDetect: Iterative Mislabel Detection using Early Loss
Summary: MisDetect identifies label noise during training by iteratively flagging high early-loss examples, applying influence-based verification, and auto-stopping when early-loss signals fade. For ambiguous instances it generates pseudo-labels to train a binary verifier; outperforms 10 baselines on 15 datasets.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13365
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 4.1905499e-05
- Overall Rank
- 11,003 | 23.53%
- DOI
-
10.14778/3648160.3648161
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 192 |
HoloClean: Holistic Data Repairs with Probabilistic Inference |
2017 |
VLDB |
0.00035692958 |
| 1,340 |
HoloDetect: Few-Shot Learning for Error Detection |
2019 |
SIGMOD |
0.00012492795 |
| 1,869 |
Interpretable Data-Based Explanations for Fairness Debugging |
2022 |
SIGMOD |
0.00010263235 |
| 2,759 |
Complaint-driven Training Data Debugging for Query 2.0 |
2020 |
SIGMOD |
8.1646193e-05 |
| 3,767 |
Cleaning Crowdsourced Labels Using Oracles for Statistical Classification |
2019 |
VLDB |
6.7748725e-05 |
| 4,103 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
6.4460899e-05 |
| 5,254 |
CDB: A Crowd-Powered Database System |
2018 |
VLDB |
5.5991922e-05 |
| 5,369 |
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach |
2016 |
SIGMOD |
5.5436995e-05 |
| 5,386 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
5.5346315e-05 |
| 7,180 |
Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning |
2023 |
VLDB |
4.8032775e-05 |
| 7,580 |
Human-in-the-loop Outlier Detection |
2020 |
SIGMOD |
4.7023767e-05 |
| 7,798 |
CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties |
2021 |
VLDB |
4.6438053e-05 |
| 9,224 |
VisClean: Interactive Cleaning for Progressive Visualization |
2020 |
VLDB |
4.3657563e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 1,340 |
HoloDetect: Few-Shot Learning for Error Detection |
2019 |
SIGMOD |
0.00012492795 |
| 11,055 |
Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines |
2024 |
VLDB |
4.1905499e-05 |
| 8,713 |
LANCET: Labeling Complex Data at Scale |
2021 |
VLDB |
4.4577046e-05 |
| 10,956 |
Certain and Approximately Certain Models for Statistical Learning |
2024 |
SIGMOD |
4.1905499e-05 |
| 10,488 |
Data Enhancement for Binary Classification of Relational Data |
2025 |
SIGMOD |
4.1905499e-05 |
| 8,588 |
Exploratory Training: When Annotators Learn About Data |
2023 |
SIGMOD |
4.4853244e-05 |
| 6,139 |
Finding Label and Model Errors in Perception Data With Learned Observation Assertions |
2022 |
SIGMOD |
5.1893488e-05 |
| 9,895 |
Towards Interpretable and Learnable Risk Analysis for Entity Resolution |
2020 |
SIGMOD |
4.2559233e-05 |
| 4,113 |
Learning to Validate the Predictions of Black Box Classifiers on Unseen Data |
2020 |
SIGMOD |
6.4326771e-05 |
| 10,537 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
4.1905499e-05 |