ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning
Summary: ActiveClean is a progressive data-cleaning framework that interleaves cleaning with ML training, updating models as analysts clean small data batches. Key ideas include importance weighting, dirty-data detection, and a visual interface, enabling robust learning in high-dimensional pipelines, demonstrated on video classification and topic modeling. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sanjay Krishnan (University of California Berkeley)
- 2. Michael J. Franklin (University of California Berkeley)
- 3. Ken Goldberg (University of California Berkeley)
- 4. Jiannan Wang (Simon Fraser University)
- 5. Eugene Wu (Columbia University)
BibTeX Citation
@inproceedings{krishnan_sigmod16,
title = {{ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning}},
author = {Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Wang, Jiannan and Wu, Eugene},
series = {{SIGMOD} '16},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2882903.2899409},
url = {https://dl.acm.org/doi/10.1145/2882903.2899409},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00017584249 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011073863 |
| 2,372 | SCODED: Statistical Constraint Oriented Data Error Detection | 2020 | SIGMOD | 8.5575299e-05 |
| 5,102 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD | 6.2705617e-05 |
| 5,956 | Semi-Supervised Data Cleaning with Raha and Baran | 2021 | CIDR | 5.9329823e-05 |
| 7,070 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.6056639e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 433 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018324965 |
| 652 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.00015114401 |
| 716 | Guided Data Repair | 2011 | VLDB | 0.00014546968 |
| 1,722 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.7921604e-05 |
| 8,785 | Wisteria: Nurturing Scalable Data Cleaning Infrastructure | 2015 | VLDB | 5.2767488e-05 |
| 8,876 | Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views | 2015 | VLDB | 5.2587627e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 204 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB |
| 2 | 11,408 | DemandClean: A Multi-Objective Learning Framework for Balancing Model Tolerance to Data Authenticity and Diversity | 2025 | VLDB |
| 3 | 12,055 | Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop | 2020 | CIDR |
| 4 | 10,984 | Minimal Data Cleaning for Model Training by MinPrep | 2026 | VLDB |
| 5 | 7,302 | CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning | 2017 | VLDB |
| 6 | 7,880 | Learning Over Dirty Data Without Cleaning | 2020 | SIGMOD |
| 7 | 1,044 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |
| 8 | 7,538 | PIClean: A Probabilistic and Interactive Data Cleaning System | 2019 | SIGMOD |
| 9 | 9,481 | VisClean: Interactive Cleaning for Progressive Visualization | 2020 | VLDB |
| 10 | 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB |