MisDetect: Iterative Mislabel Detection using Early Loss
Summary: MisDetect identifies label noise during training by iteratively flagging high early-loss examples, applying influence-based verification, and auto-stopping when early-loss signals fade. For ambiguous instances it generates pseudo-labels to train a binary verifier; outperforms 10 baselines on 15 datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yuhao Deng (Beijing Institute of Technology)
- 2. Chengliang Chai (Beijing Institute of Technology)
- 3. Lei Cao (Massachusetts Institute of Technology; University of Arizona)
- 4. Nan Tang (Hong Kong University of Science and Technology)
- 5. Jiayi Wang (Tsinghua University)
- 6. Ju Fan (Renmin University of China)
- 7. Ye Yuan (Beijing Institute of Technology)
- 8. Guoren Wang (Beijing Institute of Technology)
BibTeX Citation
@article{deng_vldb24,
title = {{MisDetect: Iterative Mislabel Detection using Early Loss}},
author = {Deng, Yuhao and Chai, Chengliang and Cao, Lei and Tang, Nan and Wang, Jiayi and Fan, Ju and Yuan, Ye and Wang, Guoren},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {6},
pages = {1159--1172},
doi = {10.14778/3648160.3648161},
url = {https://doi.org/10.14778/3648160.3648161},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,409 | LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes | 2024 | VLDB | 6.6120807e-05 |
| 7,924 | Outlier Summarization via Human Interpretable Rules | 2024 | VLDB | 5.425954e-05 |
| 11,001 | ARGO: An Interactive Data Governance System for Machine Learning via Hierarchical Reinforcement Learning | 2026 | VLDB | 4.9793485e-05 |
| 11,101 | Agree to Disagree: Robust Anomaly Detection with Noisy Labels | 2025 | SIGMOD | 4.9793485e-05 |
| 11,213 | Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection | 2025 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,589 | Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines | 2024 | VLDB |
| 2 | 9,036 | LANCET: Labeling Complex Data at Scale | 2021 | VLDB |
| 3 | 11,514 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 4 | 11,182 | Data Enhancement for Binary Classification of Relational Data | 2025 | SIGMOD |
| 5 | 8,756 | Exploratory Training: When Annotators Learn About Data | 2023 | SIGMOD |
| 6 | 10,975 | Minimal Data Cleaning for Model Training by MinPrep | 2026 | VLDB |
| 7 | 6,552 | Finding Label and Model Errors in Perception Data With Learned Observation Assertions | 2022 | SIGMOD |
| 8 | 10,244 | Towards Interpretable and Learnable Risk Analysis for Entity Resolution | 2020 | SIGMOD |
| 9 | 3,735 | Learning to Validate the Predictions of Black Box Classifiers on Unseen Data | 2020 | SIGMOD |
| 10 | 11,213 | Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection | 2025 | SIGMOD |