| 788 |
ActiveClean: Interactive Data Cleaning For Statistical Modeling |
2016 |
VLDB |
0.00016618698 |
| 1,354 |
Northstar: An Interactive Data Science System |
2018 |
VLDB |
0.00012424105 |
| 1,629 |
Data Cleaning: Overview and Emerging Challenges |
2016 |
SIGMOD |
0.00011073148 |
| 1,867 |
Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems |
2014 |
SIGMOD |
0.00010264932 |
| 1,884 |
Tuplex: Data Science in Python at Native Code Speed |
2021 |
SIGMOD |
0.00010206514 |
| 1,895 |
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning |
2020 |
VLDB |
0.00010174634 |
| 2,135 |
Towards Sustainable Insights or why polygamy is bad for you |
2017 |
CIDR |
9.4681152e-05 |
| 2,308 |
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions |
2021 |
VLDB |
9.0634287e-05 |
| 2,805 |
Query-Oriented Data Cleaning with Oracles |
2015 |
SIGMOD |
8.103731e-05 |
| 3,000 |
BigDansing: A System for Big Data Cleansing |
2015 |
SIGMOD |
7.7447724e-05 |
| 3,268 |
QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications |
2015 |
SIGMOD |
7.3027561e-05 |
| 3,767 |
Cleaning Crowdsourced Labels Using Oracles for Statistical Classification |
2019 |
VLDB |
6.7748725e-05 |
| 3,944 |
AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics |
2018 |
SIGMOD |
6.6056349e-05 |
| 4,271 |
Cleaning Denial Constraint Violations through Relaxation |
2020 |
SIGMOD |
6.2943273e-05 |
| 4,372 |
Sample Debiasing in the Themis Open World Database System |
2020 |
SIGMOD |
6.2367043e-05 |
| 4,453 |
CLAMShell: Speeding up Crowds for Low-latency Data Labeling |
2016 |
VLDB |
6.1690121e-05 |
| 4,664 |
PrivateClean: Data Cleaning and Differential Privacy |
2016 |
SIGMOD |
6.0058132e-05 |
| 5,150 |
Horizon: Scalable Dependency-driven Data Cleaning |
2021 |
VLDB |
5.6553571e-05 |
| 5,594 |
QuERy: A Framework for Integrating Entity Resolution with Query Processing |
2016 |
VLDB |
5.416945e-05 |
| 5,930 |
ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning |
2016 |
SIGMOD |
5.2632185e-05 |
| 6,690 |
Efficient Knowledge Graph Accuracy Evaluation |
2019 |
VLDB |
4.9575975e-05 |
| 6,724 |
Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing |
2021 |
SIGMOD |
4.9449472e-05 |
| 7,013 |
Qualitative Data Cleaning |
2016 |
VLDB |
4.8576683e-05 |
| 7,116 |
Crowdsourced Data Management: Overview and Challenges |
2017 |
SIGMOD |
4.8219732e-05 |
| 7,233 |
CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning |
2017 |
VLDB |
4.788267e-05 |
| 7,246 |
Learning to Sample: Counting with Complex Queries |
2020 |
VLDB |
4.7847433e-05 |
| 7,634 |
ReStore - Neural Data Completion for Relational Databases |
2021 |
SIGMOD |
4.6866388e-05 |
| 7,769 |
ICARUS: Minimizing Human Effort in Iterative Data Completion |
2018 |
VLDB |
4.6520279e-05 |
| 8,590 |
Wisteria: Nurturing Scalable Data Cleaning Infrastructure |
2015 |
VLDB |
4.4851741e-05 |
| 8,703 |
Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views |
2015 |
VLDB |
4.4596255e-05 |
| 9,043 |
Query-Guided Resolution in Uncertain Databases |
2023 |
SIGMOD |
4.3997447e-05 |
| 9,053 |
Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise |
2019 |
VLDB |
4.3997447e-05 |
| 9,055 |
A Data Quality Metric (DQM): How to Estimate the Number of Undetected Errors in Data Sets |
2017 |
VLDB |
4.3997447e-05 |
| 9,218 |
QOCO: A Query Oriented Data Cleaning System with Oracles |
2015 |
VLDB |
4.3672293e-05 |
| 9,354 |
GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models |
2024 |
SIGMOD |
4.3484715e-05 |
| 10,625 |
Deduplicated Sampling On-Demand |
2025 |
VLDB |
4.1905499e-05 |
| 11,032 |
Efficient and Reliable Estimation of Knowledge Graph Accuracy |
2024 |
VLDB |
4.1905499e-05 |