DBScholar

Back to papers

ActiveClean: Interactive Data Cleaning For Statistical Modeling

Summary: ActiveClean enables progressive, iterative cleaning during convex-loss model training while preserving convergence guarantees. It prioritizes records most likely to affect model parameters, achieving substantially higher accuracy than uniform sampling and active learning under fixed cleaning budgets. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h8d50740df96ad7e5
Venue
VLDB
Year
2016
Pagerank
0.00017590977
Overall Rank
483 | 96.76%
DOI
10.14778/2994509.2994511

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{krishnan_vldb16,
        title = {{ActiveClean: Interactive Data Cleaning For Statistical Modeling}},
        author = {Krishnan, Sanjay and Wang, Jiannan and Wu, Eugene and Franklin, Michael J. and Goldberg, Ken},
        journal = {PVLDB},
        series = {{VLDB} '16},
        volume = {9},
        number = {12},
        pages = {948--959},
        doi = {10.14778/2994509.2994511},
        url = {https://doi.org/10.14778/2994509.2994511},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 53 citing papers.

Rank Citing Paper Year Venue Pagerank
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,255 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011325762
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010968223
1,805 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 9.59842e-05
1,847 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5120573e-05
2,598 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.2385793e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8742664e-05
2,960 Auto-Detect: Data-Driven Error Detection in Tables 2018 SIGMOD 7.8075103e-05
3,263 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4809771e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
3,331 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 7.4166095e-05
3,360 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 7.3767115e-05
3,941 GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data 2023 SIGMOD 6.9138042e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,195 PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models 2020 SIGMOD 6.7422305e-05
4,524 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.5621504e-05
4,719 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 6.4548076e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
5,247 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 6.2121767e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
6,115 Equitable Data Valuation Meets the Right to Be Forgotten in Model Markets 2023 VLDB 5.8817545e-05
6,921 CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning 2024 SIGMOD 5.6432616e-05
7,538 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.4995874e-05
7,875 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.4355826e-05
7,876 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 5.4348457e-05
8,411 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.336291e-05
8,756 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2840289e-05
9,004 The Cost of Representation by Subset Repairs 2025 VLDB 5.2389669e-05
9,373 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.1868213e-05
9,384 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.1868213e-05
9,417 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 5.1803615e-05
9,565 Deduplicated Sampling On-Demand 2025 VLDB 5.1571823e-05
9,616 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1510548e-05
9,722 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1349531e-05
10,515 Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] 2026 SIGMOD 4.9793485e-05
10,529 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.9793485e-05
10,530 Outliers: The Good, the Bad and the Ugly 2026 SIGMOD 4.9793485e-05
10,799 Finding Non-Redundant Simpson's Paradox in Multidimensional Data 2026 VLDB 4.9793485e-05
10,855 PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines 2026 VLDB 4.9793485e-05
10,975 Minimal Data Cleaning for Model Training by MinPrep 2026 VLDB 4.9793485e-05
10,981 NiceT: Named Entity Cleaning and Enhancement with Human-in-the-loop 2026 VLDB 4.9793485e-05
11,001 ARGO: An Interactive Data Governance System for Machine Learning via Hierarchical Reinforcement Learning 2026 VLDB 4.9793485e-05
11,182 Data Enhancement for Binary Classification of Relational Data 2025 SIGMOD 4.9793485e-05
11,213 Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection 2025 SIGMOD 4.9793485e-05
11,284 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9793485e-05
11,294 Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation 2025 VLDB 4.9793485e-05
11,407 mlidea: Interactively Improving ML Data Preparation Code via “Shadow Pipelines” 2025 VLDB 4.9793485e-05
11,514 Certain and Approximately Certain Models for Statistical Learning 2024 SIGMOD 4.9793485e-05
11,589 Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines 2024 VLDB 4.9793485e-05
11,661 Generalizable Data Cleaning of Tabular Data in Latent Space 2024 VLDB 4.9793485e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers