| 740 |
Distributed Representations of Tuples for Entity Resolution |
2018 |
VLDB |
0.00017358024 |
| 1,012 |
NADEEF: A Commodity Data Cleaning System |
2013 |
SIGMOD |
0.00014638349 |
| 1,160 |
Towards Certain Fixes with Editing Rules and Master Data |
2010 |
VLDB |
0.0001358129 |
| 1,274 |
The Data Civilizer System |
2017 |
CIDR |
0.00012869297 |
| 1,403 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00012180046 |
| 1,419 |
Graph Pattern Matching: From Intractable to Polynomial Time |
2010 |
VLDB |
0.00012072488 |
| 1,505 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00011601232 |
| 1,544 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
0.00011438274 |
| 1,821 |
Synthesizing Entity Matching Rules by Examples |
2018 |
VLDB |
0.00010406856 |
| 1,893 |
Querying Shortest Paths on Time Dependent Road Networks |
2019 |
VLDB |
0.00010175836 |
| 2,348 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
8.9903659e-05 |
| 2,609 |
Graph Stream Summarization: From Big Bang to Big Crunch |
2016 |
SIGMOD |
8.4587236e-05 |
| 2,830 |
Interaction between Record Matching and Data Repairing |
2011 |
SIGMOD |
8.0515409e-05 |
| 2,948 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
7.8301799e-05 |
| 2,968 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
7.7964476e-05 |
| 3,000 |
BigDansing: A System for Big Data Cleansing |
2015 |
SIGMOD |
7.7447724e-05 |
| 3,198 |
Towards Dependable Data Repairing with Fixing Rules |
2014 |
SIGMOD |
7.4029546e-05 |
| 3,270 |
RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - |
2018 |
VLDB |
7.3013196e-05 |
| 3,455 |
Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation |
2022 |
VLDB |
7.0760196e-05 |
| 3,469 |
Deep Learning for Blocking in Entity Matching: A Design Space Exploration |
2021 |
VLDB |
7.0629476e-05 |
| 3,568 |
NADEEF/ER: Generic and Interactive Entity Resolution |
2014 |
SIGMOD |
6.9588883e-05 |
| 3,575 |
Lightning Fast and Space Efficient Inequality Joins |
2015 |
VLDB |
6.9509846e-05 |
| 3,666 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
6.8606092e-05 |
| 3,854 |
Generating Concise Entity Matching Rules |
2017 |
SIGMOD |
6.697423e-05 |
| 3,973 |
HAIChart: Human and AI Paired Visualization System |
2024 |
VLDB |
6.5721521e-05 |
| 3,977 |
UGuide – User-Guided Discovery of FD-Detectable Errors |
2017 |
SIGMOD |
6.5670094e-05 |
| 4,103 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
6.4460899e-05 |
| 4,211 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
6.3495931e-05 |
| 4,829 |
Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks |
2021 |
SIGMOD |
5.8890126e-05 |
| 4,908 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
5.835596e-05 |
| 5,026 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
5.7451454e-05 |
| 5,056 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
5.7225065e-05 |
| 5,203 |
Pattern Functional Dependencies for Data Cleaning |
2020 |
VLDB |
5.628087e-05 |
| 5,207 |
ANMAT: Automatic Knowledge Discovery and Error Detection through Pattern Functional Dependencies |
2019 |
SIGMOD |
5.6254804e-05 |
| 5,386 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
5.5346315e-05 |
| 5,470 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
5.4894925e-05 |
| 5,477 |
Learned Cardinality Estimation for Similarity Queries |
2021 |
SIGMOD |
5.4856699e-05 |
| 5,493 |
DeepEye: Creating Good Data Visualizations by Keyword Search |
2018 |
SIGMOD |
5.4773922e-05 |
| 5,698 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
5.3669165e-05 |
| 5,738 |
KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing |
2015 |
VLDB |
5.3454984e-05 |
| 5,965 |
Automatic Data Acquisition for Deep Learning |
2021 |
VLDB |
5.2476363e-05 |
| 6,279 |
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks |
2023 |
VLDB |
5.1241232e-05 |
| 6,348 |
NADEEF: A Generalized Data Cleaning System |
2013 |
VLDB |
5.0969173e-05 |
| 6,564 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
5.0037892e-05 |
| 6,570 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
5.0017341e-05 |
| 6,840 |
Towards Democratizing Relational Data Visualization |
2019 |
SIGMOD |
4.9058387e-05 |
| 7,180 |
Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning |
2023 |
VLDB |
4.8032775e-05 |
| 7,586 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
4.700127e-05 |
| 7,992 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
4.6082565e-05 |
| 8,121 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
4.5771143e-05 |