| 7,068 |
How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses |
2024 |
VLDB |
5.6083188e-05 |
| 7,183 |
DataPrism: Exposing Disconnect between Data and Systems |
2022 |
SIGMOD |
5.5917354e-05 |
| 7,250 |
Akane: Perplexity-Guided Time Series Data Cleaning |
2024 |
SIGMOD |
5.574104e-05 |
| 7,291 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
5.5642829e-05 |
| 7,415 |
Data Integration and Machine Learning: A Natural Synergy |
2018 |
VLDB |
5.5332662e-05 |
| 7,471 |
Fast Detection of Denial Constraint Violations |
2022 |
VLDB |
5.5176505e-05 |
| 7,532 |
PIClean: A Probabilistic and Interactive Data Cleaning System |
2019 |
SIGMOD |
5.5007996e-05 |
| 7,583 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
5.4906208e-05 |
| 7,648 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5.4772833e-05 |
| 7,704 |
ReStore - Neural Data Completion for Relational Databases |
2021 |
SIGMOD |
5.4741304e-05 |
| 7,839 |
ExDRa: Exploratory Data Science on Federated Raw Data |
2021 |
SIGMOD |
5.4432099e-05 |
| 7,850 |
CoClean: Collaborative Data Cleaning |
2020 |
SIGMOD |
5.4405793e-05 |
| 7,872 |
Rock: Cleaning Data by Embedding ML in Logic Rules |
2024 |
SIGMOD |
5.4362062e-05 |
| 7,875 |
Learning Over Dirty Data Without Cleaning |
2020 |
SIGMOD |
5.4355826e-05 |
| 8,062 |
Evaluating Top-k Queries with Inconsistency Degrees |
2020 |
VLDB |
5.3955302e-05 |
| 8,224 |
Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness |
2024 |
VLDB |
5.3747673e-05 |
| 8,277 |
Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale |
2022 |
VLDB |
5.3639084e-05 |
| 8,280 |
The Computation of Optimal Subset Repairs |
2020 |
VLDB |
5.363617e-05 |
| 8,335 |
ICARUS: Minimizing Human Effort in Iterative Data Completion |
2018 |
VLDB |
5.3525933e-05 |
| 8,398 |
Deducing Certain Fixes to Graphs |
2019 |
VLDB |
5.3399392e-05 |
| 8,411 |
SHiFT: An Efficient, Flexible Search Engine for Transfer Learning |
2023 |
VLDB |
5.336291e-05 |
| 8,544 |
Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? |
2021 |
SIGMOD |
5.3186267e-05 |
| 8,695 |
From Papers to Practice: The openclean Open-Source Data Cleaning Library |
2021 |
VLDB |
5.2905577e-05 |
| 8,756 |
Exploratory Training: When Annotators Learn About Data |
2023 |
SIGMOD |
5.2840289e-05 |
| 8,847 |
Machine Learning Meets Big Spatial Data |
2019 |
VLDB |
5.2646682e-05 |
| 8,852 |
Rapidash: Efficient Detection of Constraint Violations |
2024 |
VLDB |
5.2641151e-05 |
| 8,886 |
FastPDB: Towards Bag-Probabilistic Queries at Interactive Speeds |
2025 |
SIGMOD |
5.2559789e-05 |
| 8,984 |
nsDB: Architecting the Next Generation Database by Integrating Neural and Symbolic Systems |
2024 |
VLDB |
5.2433452e-05 |
| 9,004 |
The Cost of Representation by Subset Repairs |
2025 |
VLDB |
5.2389669e-05 |
| 9,229 |
Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables |
2025 |
SIGMOD |
5.2056825e-05 |
| 9,294 |
DataDiff: User-Interpretable Data Transformation Summaries for Collaborative Data Analysis |
2018 |
SIGMOD |
5.1995938e-05 |
| 9,311 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
5.1965878e-05 |
| 9,373 |
Query-Guided Resolution in Uncertain Databases |
2023 |
SIGMOD |
5.1868213e-05 |
| 9,384 |
Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise |
2019 |
VLDB |
5.1868213e-05 |
| 9,417 |
Towards Observability for Production Machine Learning Pipelines |
2022 |
VLDB |
5.1803615e-05 |
| 9,455 |
ZIP: Lazy Imputation during Query Processing |
2024 |
VLDB |
5.1735482e-05 |
| 9,608 |
Discovering Top-k Rules using Subjective and Objective Criteria |
2023 |
SIGMOD |
5.1527671e-05 |
| 9,616 |
GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models |
2024 |
SIGMOD |
5.1510548e-05 |
| 9,722 |
DataVinci: Learning Syntactic and Semantic String Repairs |
2025 |
SIGMOD |
5.1349531e-05 |
| 9,766 |
Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation |
2021 |
CIDR |
5.1344318e-05 |
| 9,801 |
Incremental Detection of Denial Constraint Violations |
2025 |
VLDB |
5.1257999e-05 |
| 9,806 |
Making It Tractable to Catch Duplicates and Conflicts in Graphs |
2023 |
SIGMOD |
5.1257999e-05 |
| 9,808 |
Lingua Manga: A Generic Large Language Model Centric System for Data Curation |
2023 |
VLDB |
5.1257999e-05 |
| 9,874 |
MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data |
2024 |
VLDB |
5.1176637e-05 |
| 10,025 |
Don’t Be a Tattle-Tale: Preventing Leakages through Data Dependencies on Access Control Protected Data |
2022 |
VLDB |
5.0934585e-05 |
| 10,106 |
Efficient Differential Dependency Discovery |
2024 |
VLDB |
5.0789354e-05 |
| 10,110 |
EasyDR: A Human-in-the-loop Error Detection&Repair Platform for Holistic Table Cleaning |
2022 |
VLDB |
5.0789354e-05 |
| 10,181 |
In-Database Data Imputation |
2024 |
SIGMOD |
5.0653015e-05 |
| 10,186 |
Discovering Top-k Relevant and Diversified Rules |
2024 |
SIGMOD |
5.0651993e-05 |
| 10,188 |
Reptile: Aggregation-level Explanations for Hierarchical Data |
2022 |
SIGMOD |
5.0651993e-05 |