ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning
Summary: ActiveClean is a progressive data-cleaning framework that interleaves cleaning with ML training, updating models as analysts clean small data batches. Key ideas include importance weighting, dirty-data detection, and a visual interface, enabling robust learning in high-dimensional pipelines, demonstrated on video classification and topic modeling. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sanjay Krishnan (University of California Berkeley)
- 2. Michael J. Franklin (University of California Berkeley)
- 3. Ken Goldberg (University of California Berkeley)
- 4. Jiannan Wang (Simon Fraser University)
- 5. Eugene Wu (Columbia University)
BibTeX Citation
@inproceedings{krishnan_sigmod16,
title = {{ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning}},
author = {Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Wang, Jiannan and Wu, Eugene},
series = {{SIGMOD} '16},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2882903.2899409},
url = {https://dl.acm.org/doi/10.1145/2882903.2899409},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB | 0.00016148948 |
| 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.00011065626 |
| 3,004 | SCODED: Statistical Constraint Oriented Data Error Detection | 2020 | SIGMOD | 7.8608629e-05 |
| 5,051 | Towards Benchmarking Feature Type Inference for AutoML Platforms | 2021 | SIGMOD | 6.385354e-05 |
| 5,835 | Semi-Supervised Data Cleaning with Raha and Baran | 2021 | CIDR | 6.0720263e-05 |
| 6,928 | How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses | 2024 | VLDB | 5.7370426e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 439 | Corleone: Hands-Off Crowdsourcing for Entity Matching | 2014 | SIGMOD | 0.00018464913 |
| 661 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.0001519162 |
| 714 | Guided Data Repair | 2011 | VLDB | 0.00014662041 |
| 1,736 | A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data | 2014 | SIGMOD | 9.8984415e-05 |
| 8,620 | Wisteria: Nurturing Scalable Data Cleaning Infrastructure | 2015 | VLDB | 5.3991092e-05 |
| 8,714 | Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views | 2015 | VLDB | 5.3778009e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 201 | Declarative Data Cleaning: Language, Model, and Algorithms | 2001 | VLDB |
| 2 | 13,435 | Data Cleaning in the Era of Data Science: Challenges and Opportunities | 2021 | CIDR |
| 3 | 11,038 | DemandClean: A Multi-Objective Learning Framework for Balancing Model Tolerance to Data Authenticity and Diversity | 2025 | VLDB |
| 4 | 11,746 | Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop | 2020 | CIDR |
| 5 | 7,179 | CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning | 2017 | VLDB |
| 6 | 7,880 | Learning Over Dirty Data Without Cleaning | 2020 | SIGMOD |
| 7 | 1,323 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |
| 8 | 7,387 | PIClean: A Probabilistic and Interactive Data Cleaning System | 2019 | SIGMOD |
| 9 | 9,298 | VisClean: Interactive Cleaning for Progressive Visualization | 2020 | VLDB |
| 10 | 582 | ActiveClean: Interactive Data Cleaning For Statistical Modeling | 2016 | VLDB |