| 1,422 |
Data Management Challenges in Production Machine Learning |
2017 |
SIGMOD |
0.00012050431 |
| 1,534 |
Data Management in Machine Learning: Challenges, Techniques, and Systems |
2017 |
SIGMOD |
0.00011462072 |
| 1,895 |
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning |
2020 |
VLDB |
0.00010174634 |
| 2,308 |
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions |
2021 |
VLDB |
9.0634287e-05 |
| 2,507 |
Auto-Detect: Data-Driven Error Detection in Tables |
2018 |
SIGMOD |
8.6254741e-05 |
| 2,759 |
Complaint-driven Training Data Debugging for Query 2.0 |
2020 |
SIGMOD |
8.1646193e-05 |
| 2,845 |
VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space Decomposition |
2021 |
VLDB |
8.0301674e-05 |
| 2,968 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
7.7964476e-05 |
| 3,397 |
Automatic Data Repair: Are We Ready to Deploy? |
2024 |
VLDB |
7.1386386e-05 |
| 3,466 |
AI Meets Database: AI4DB and DB4AI |
2021 |
SIGMOD |
7.0645718e-05 |
| 3,767 |
Cleaning Crowdsourced Labels Using Oracles for Statistical Classification |
2019 |
VLDB |
6.7748725e-05 |
| 4,103 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
6.4460899e-05 |
| 4,271 |
Cleaning Denial Constraint Violations through Relaxation |
2020 |
SIGMOD |
6.2943273e-05 |
| 4,423 |
PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models |
2020 |
SIGMOD |
6.1925724e-05 |
| 4,596 |
Data Integration and Machine Learning: A Natural Synergy |
2018 |
SIGMOD |
6.0540725e-05 |
| 4,866 |
OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine Learning |
2021 |
SIGMOD |
5.8620848e-05 |
| 5,227 |
Enabling SQL-based Training Data Debugging for Federated Learning |
2022 |
VLDB |
5.6156523e-05 |
| 5,439 |
DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data |
2023 |
SIGMOD |
5.5034427e-05 |
| 5,974 |
Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond |
2021 |
SIGMOD |
5.2458154e-05 |
| 6,263 |
Equitable Data Valuation Meets the Right to Be Forgotten in Model Markets |
2023 |
VLDB |
5.1300221e-05 |
| 7,798 |
CHEF: A Cheap and Fast Pipeline for Iteratively Cleaning Label Uncertainties |
2021 |
VLDB |
4.6438053e-05 |
| 7,868 |
Learning Over Dirty Data Without Cleaning |
2020 |
SIGMOD |
4.6276013e-05 |
| 8,096 |
Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications |
2023 |
SIGMOD |
4.583522e-05 |
| 8,183 |
SHiFT: An Efficient, Flexible Search Engine for Transfer Learning |
2023 |
VLDB |
4.5615358e-05 |
| 8,253 |
Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines |
2023 |
SIGMOD |
4.5444167e-05 |
| 8,588 |
Exploratory Training: When Annotators Learn About Data |
2023 |
SIGMOD |
4.4853244e-05 |
| 8,739 |
CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning |
2024 |
SIGMOD |
4.4520434e-05 |
| 8,839 |
The Cost of Representation by Subset Repairs |
2025 |
VLDB |
4.4346105e-05 |
| 9,043 |
Query-Guided Resolution in Uncertain Databases |
2023 |
SIGMOD |
4.3997447e-05 |
| 9,053 |
Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise |
2019 |
VLDB |
4.3997447e-05 |
| 9,116 |
Towards Observability for Production Machine Learning Pipelines |
2022 |
VLDB |
4.3886184e-05 |
| 9,354 |
GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models |
2024 |
SIGMOD |
4.3484715e-05 |
| 9,395 |
DataVinci: Learning Syntactic and Semantic String Repairs |
2025 |
SIGMOD |
4.3399748e-05 |
| 10,026 |
Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,029 |
Outliers: The Good, the Bad and the Ugly |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,488 |
Data Enhancement for Binary Classification of Relational Data |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,537 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,625 |
Deduplicated Sampling On-Demand |
2025 |
VLDB |
4.1905499e-05 |
| 10,636 |
CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines |
2025 |
VLDB |
4.1905499e-05 |
| 10,652 |
Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation |
2025 |
VLDB |
4.1905499e-05 |
| 10,821 |
mlidea: Interactively Improving ML Data Preparation Code via "Shadow Pipelines" |
2025 |
VLDB |
4.1905499e-05 |
| 10,956 |
Certain and Approximately Certain Models for Statistical Learning |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,055 |
Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines |
2024 |
VLDB |
4.1905499e-05 |
| 11,140 |
Generalizable Data Cleaning of Tabular Data in Latent Space |
2024 |
VLDB |
4.1905499e-05 |
| 11,181 |
LinCQA: Faster Consistent Query Answering with Linear Time Guarantees |
2023 |
SIGMOD |
4.1905499e-05 |
| 11,434 |
Ease.ML: A Lifecycle Management System for MLDev and MLOps |
2021 |
CIDR |
4.1905499e-05 |
| 11,687 |
IHCS: An Integrated Hybrid Cleaning System |
2019 |
VLDB |
4.1905499e-05 |