MisDetect: Iterative Mislabel Detection using Early Loss
Summary: MisDetect identifies label noise during training by iteratively flagging high early-loss examples, applying influence-based verification, and auto-stopping when early-loss signals fade. For ambiguous instances it generates pseudo-labels to train a binary verifier; outperforms 10 baselines on 15 datasets. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yuhao Deng (Beijing Institute of Technology)
- 2. Chengliang Chai (Beijing Institute of Technology)
- 3. Lei Cao (Massachusetts Institute of Technology; University of Arizona)
- 4. Nan Tang (Hong Kong University of Science and Technology)
- 5. Jiayi Wang (Tsinghua University)
- 6. Ju Fan (Renmin University of China)
- 7. Ye Yuan (Beijing Institute of Technology)
- 8. Guoren Wang (Beijing Institute of Technology)
BibTeX Citation
@article{deng_vldb24,
title = {{MisDetect: Iterative Mislabel Detection using Early Loss}},
author = {Deng, Yuhao and Chai, Chengliang and Cao, Lei and Tang, Nan and Wang, Jiayi and Fan, Ju and Yuan, Ye and Wang, Guoren},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {6},
pages = {1159--1172},
doi = {10.14778/3648160.3648161},
url = {https://doi.org/10.14778/3648160.3648161},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,893 | LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes | 2024 | VLDB | 5.5209866e-05 |
| 8,802 | Outlier Summarization via Human Interpretable Rules | 2024 | VLDB | 5.3685765e-05 |
| 10,658 | Agree to Disagree: Robust Anomaly Detection with Noisy Labels | 2025 | SIGMOD | 5.093636e-05 |
| 10,800 | Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection | 2025 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 946 | HoloDetect: Few-Shot Learning for Error Detection | 2019 | SIGMOD |
| 2 | 11,260 | Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines | 2024 | VLDB |
| 3 | 8,876 | LANCET: Labeling Complex Data at Scale | 2021 | VLDB |
| 4 | 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD |
| 5 | 10,757 | Data Enhancement for Binary Classification of Relational Data | 2025 | SIGMOD |
| 6 | 8,594 | Exploratory Training: When Annotators Learn About Data | 2023 | SIGMOD |
| 7 | 6,431 | Finding Label and Model Errors in Perception Data With Learned Observation Assertions | 2022 | SIGMOD |
| 8 | 10,049 | Towards Interpretable and Learnable Risk Analysis for Entity Resolution | 2020 | SIGMOD |
| 9 | 3,670 | Learning to Validate the Predictions of Black Box Classifiers on Unseen Data | 2020 | SIGMOD |
| 10 | 10,800 | Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection | 2025 | SIGMOD |