DBScholar

Back to papers

ActiveClean: Interactive Data Cleaning For Statistical Modeling

Summary: ActiveClean enables progressive, iterative cleaning during convex-loss model training while preserving convergence guarantees. It prioritizes records most likely to affect model parameters, achieving substantially higher accuracy than uniform sampling and active learning under fixed cleaning budgets. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h8d50740df96ad7e5
Venue
VLDB
Year
2016
Pagerank
0.00017584249
Overall Rank
483 | 96.76%
DOI
10.14778/2994509.2994511
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{krishnan_vldb16,
        title = {{ActiveClean: Interactive Data Cleaning For Statistical Modeling}},
        author = {Krishnan, Sanjay and Wang, Jiannan and Wu, Eugene and Franklin, Michael J. and Goldberg, Ken},
        journal = {PVLDB},
        series = {{VLDB} '16},
        volume = {9},
        number = {12},
        pages = {948--959},
        doi = {10.14778/2994509.2994511},
        url = {https://doi.org/10.14778/2994509.2994511},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 53 citing papers.

Rank Citing Paper Year Venue Pagerank
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011793347
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011468426
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010963254
1,805 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 9.5938877e-05
1,848 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5075544e-05
2,600 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.2346824e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8716173e-05
2,962 Auto-Detect: Data-Driven Error Detection in Tables 2018 SIGMOD 7.8038782e-05
3,264 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4774357e-05
3,308 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4379101e-05
3,332 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.4130985e-05
3,360 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 7.3732195e-05
3,942 GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data 2023 SIGMOD 6.9105312e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7576159e-05
4,197 PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models 2020 SIGMOD 6.7390419e-05
4,525 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.559044e-05
4,721 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.451752e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,251 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 6.2092359e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
6,116 Equitable Data Valuation Meets the Right to Be Forgotten in Model Markets 2023 VLDB 5.8789701e-05
6,924 CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning 2024 SIGMOD 5.6405901e-05
7,544 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.496984e-05
7,880 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.4330096e-05
7,881 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4322729e-05
8,418 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.3337649e-05
8,764 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2815275e-05
9,014 The Cost of Representation by Subset Repairs 2025 VLDB 5.2364868e-05
9,382 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.1843659e-05
9,393 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.1843659e-05
9,426 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 5.1779092e-05
9,573 Deduplicated Sampling On-Demand 2025 VLDB 5.154741e-05
9,623 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1486163e-05
9,727 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1325223e-05
10,526 Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] 2026 SIGMOD 4.9769913e-05
10,540 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.9769913e-05
10,541 Outliers: The Good, the Bad and the Ugly 2026 SIGMOD 4.9769913e-05
10,809 Finding Non-Redundant Simpson's Paradox in Multidimensional Data 2026 VLDB 4.9769913e-05
10,864 PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines 2026 VLDB 4.9769913e-05
10,984 Minimal Data Cleaning for Model Training by MinPrep 2026 VLDB 4.9769913e-05
10,990 NiceT: Named Entity Cleaning and Enhancement with Human-in-the-loop 2026 VLDB 4.9769913e-05
11,010 ARGO: An Interactive Data Governance System for Machine Learning via Hierarchical Reinforcement Learning 2026 VLDB 4.9769913e-05
11,191 Data Enhancement for Binary Classification of Relational Data 2025 SIGMOD 4.9769913e-05
11,222 Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection 2025 SIGMOD 4.9769913e-05
11,292 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9769913e-05
11,302 Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation 2025 VLDB 4.9769913e-05
11,413 mlidea: Interactively Improving ML Data Preparation Code via “Shadow Pipelines” 2025 VLDB 4.9769913e-05
11,520 Certain and Approximately Certain Models for Statistical Learning 2024 SIGMOD 4.9769913e-05
11,595 Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines 2024 VLDB 4.9769913e-05
11,667 Generalizable Data Cleaning of Tabular Data in Latent Space 2024 VLDB 4.9769913e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers