ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning
Summary: ActiveClean is a progressive data-cleaning framework that interleaves cleaning with ML training, updating models as analysts clean small data batches. Key ideas include importance weighting, dirty-data detection, and a visual interface, enabling robust learning in high-dimensional pipelines, demonstrated on video classification and topic modeling. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sanjay Krishnan (University of California Berkeley)
- 2. Michael J. Franklin (University of California Berkeley)
- 3. Ken Goldberg (University of California Berkeley)
- 4. Jiannan Wang (Simon Fraser University)
- 5. Eugene Wu (Columbia University)
BibTeX Citation
@inproceedings{krishnan_sigmod16,
title = {{ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning}},
author = {Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Wang, Jiannan and Wu, Eugene},
series = {{SIGMOD} '16},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2882903.2899409},
url = {https://dl.acm.org/doi/10.1145/2882903.2899409},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00017590977 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.0001107886 |
| 2,371 | SCODED: Statistical Constraint Oriented Data Error Detection | 2020 | SIGMOD | 8.5615698e-05 |
| 5,100 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD | 6.27349e-05 |
| 5,954 | Semi-Supervised Data Cleaning with Raha and Baran | 2021 | CIDR | 5.9357922e-05 |
| 7,068 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.6083188e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 433 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018332741 |
| 652 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.00015121325 |
| 716 | Guided Data Repair | 2011 | VLDB | 0.00014553463 |
| 1,720 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.7965659e-05 |
| 8,777 | Wisteria: Nurturing Scalable Data Cleaning Infrastructure | 2015 | VLDB | 5.2792451e-05 |
| 8,870 | Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views | 2015 | VLDB | 5.2601766e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 204 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB |
| 2 | 11,402 | DemandClean: A Multi-Objective Learning Framework for Balancing Model Tolerance to Data Authenticity and Diversity | 2025 | VLDB |
| 3 | 12,049 | Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop | 2020 | CIDR |
| 4 | 10,975 | Minimal Data Cleaning for Model Training by MinPrep | 2026 | VLDB |
| 5 | 7,298 | CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning | 2017 | VLDB |
| 6 | 7,875 | Learning Over Dirty Data Without Cleaning | 2020 | SIGMOD |
| 7 | 1,043 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |
| 8 | 7,532 | PIClean: A Probabilistic and Interactive Data Cleaning System | 2019 | SIGMOD |
| 9 | 9,470 | VisClean: Interactive Cleaning for Progressive Visualization | 2020 | VLDB |
| 10 | 483 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB |