Learning Over Dirty Data Without Cleaning
Summary: DLearn learns directly from dirty data without cleaning, bypassing data-repair bottlenecks. It leverages database constraints to infer relational models that summarize patterns across all plausible clean versions; empirical evaluation on large real-world datasets shows accuracy and efficiency. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jose Picado (Oregon State University)
- 2. John Davis (Oregon State University)
- 3. Arash Termehchy (Oregon State University)
- 4. Ga Young Lee (Oregon State University)
BibTeX Citation
@inproceedings{picado_sigmod20,
title = {{Learning Over Dirty Data Without Cleaning}},
author = {Picado, Jose and Davis, John and Termehchy, Arash and Lee, Ga Young},
series = {{SIGMOD} '20},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3318464.3389708},
url = {https://dl.acm.org/doi/10.1145/3318464.3389708},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,735 | Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation | 2022 | SIGMOD | 6.5315782e-05 |
| 9,437 | GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models | 2024 | SIGMOD | 5.2687567e-05 |
| 9,993 | In-Database Data Imputation | 2024 | SIGMOD | 5.1815618e-05 |
| 11,169 | Certain and Approximately Certain Models for Statistical Learning | 2024 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,837 | PrivateClean: Data Cleaning and Differential Privacy | 2016 | SIGMOD |
| 2 | 533 | Improving Data Quality: Consistency and Accuracy | 2007 | VLDB |
| 3 | 7,179 | CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning | 2017 | VLDB |
| 4 | 3,580 | Automatic Data Repair: Are We Ready to Deploy? | 2024 | VLDB |
| 5 | 13,435 | Data Cleaning in the Era of Data Science: Challenges and Opportunities | 2021 | CIDR |
| 6 | 647 | Discovering Data Quality Rules | 2008 | VLDB |
| 7 | 5,794 | ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning | 2016 | SIGMOD |
| 8 | 11,343 | Generalizable Data Cleaning of Tabular Data in Latent Space | 2024 | VLDB |
| 9 | 10,785 | Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables | 2025 | SIGMOD |
| 10 | 1,323 | Data Cleaning: Overview and Emerging Challenges | 2016 | SIGMOD |