Database Paper Browser

Back to papers

ActiveClean: Interactive Data Cleaning For Statistical Modeling

Summary: ActiveClean enables iterative, interactive data cleaning during model training. Convex loss models are targeted; preserves convergence and prioritizes influential records; yields up to 2.5x accuracy per cleaning budget, beating uniform sampling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11382
Venue
VLDB
Year
2016
Pagerank
0.00016618698
Overall Rank
788 | 94.53%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 47 of 47 citing papers.

Rank Citing Paper Year Venue Pagerank
1,422 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00012050431
1,534 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011462072
1,895 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010174634
2,308 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.0634287e-05
2,507 Auto-Detect: Data-Driven Error Detection in Tables 2018 SIGMOD 8.6254741e-05
2,759 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.1646193e-05
2,845 VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition 2021 VLDB 8.0301674e-05
2,968 Raha: A Configuration-Free Error Detection System 2019 SIGMOD 7.7964476e-05
3,397 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.1386386e-05
3,466 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.0645718e-05
3,767 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 6.7748725e-05
4,103 GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data 2023 SIGMOD 6.4460899e-05
4,271 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 6.2943273e-05
4,423 PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models 2020 SIGMOD 6.1925724e-05
4,596 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.0540725e-05
4,866 OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning 2021 SIGMOD 5.8620848e-05
5,227 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 5.6156523e-05
5,439 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 5.5034427e-05
5,974 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 5.2458154e-05
6,263 Equitable Data Valuation Meets the Right to Be Forgotten in Model Markets 2023 VLDB 5.1300221e-05
7,798 CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties 2021 VLDB 4.6438053e-05
7,868 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 4.6276013e-05
8,096 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 4.583522e-05
8,183 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 4.5615358e-05
8,253 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 4.5444167e-05
8,588 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 4.4853244e-05
8,739 CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning 2024 SIGMOD 4.4520434e-05
8,839 The Cost of Representation by Subset Repairs 2025 VLDB 4.4346105e-05
9,043 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 4.3997447e-05
9,053 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 4.3997447e-05
9,116 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 4.3886184e-05
9,354 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 4.3484715e-05
9,395 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 4.3399748e-05
10,026 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.1905499e-05
10,029 Outliers: The Good, the Bad and the Ugly 2026 SIGMOD 4.1905499e-05
10,488 Data Enhancement for Binary Classification of Relational Data 2025 SIGMOD 4.1905499e-05
10,537 Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection 2025 SIGMOD 4.1905499e-05
10,625 Deduplicated Sampling On-Demand 2025 VLDB 4.1905499e-05
10,636 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.1905499e-05
10,652 Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation 2025 VLDB 4.1905499e-05
10,821 mlidea: Interactively Improving ML Data Preparation Code via "Shadow Pipelines" 2025 VLDB 4.1905499e-05
10,956 Certain and Approximately Certain Models for Statistical Learning 2024 SIGMOD 4.1905499e-05
11,055 Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines 2024 VLDB 4.1905499e-05
11,140 Generalizable Data Cleaning of Tabular Data in Latent Space 2024 VLDB 4.1905499e-05
11,181 LinCQA: Faster Consistent Query Answering with Linear Time Guarantees 2023 SIGMOD 4.1905499e-05
11,434 Ease.ML: A Lifecycle Management System for MLDev and MLOps 2021 CIDR 4.1905499e-05
11,687 IHCS: An Integrated Hybrid Cleaning System 2019 VLDB 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers