| 457 |
Distributed Representations of Tuples for Entity Resolution |
2018 |
VLDB |
65 |
0.00017899824 |
| 697 |
NADEEF: A Commodity Data Cleaning System |
2013 |
SIGMOD |
58 |
0.00014687805 |
| 972 |
The Data Civilizer System |
2017 |
CIDR |
56 |
0.00012757732 |
| 989 |
Towards Certain Fixes with Editing Rules and Master Data |
2010 |
VLDB |
37 |
0.00012647631 |
| 1,098 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
48 |
0.00012031983 |
| 1,150 |
Graph Pattern Matching: From Intractable to Polynomial Time |
2010 |
VLDB |
34 |
0.00011804185 |
| 1,344 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
50 |
0.00010951939 |
| 1,387 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
19 |
0.00010829741 |
| 1,493 |
Synthesizing Entity Matching Rules by Examples |
2018 |
VLDB |
31 |
0.00010500948 |
| 1,724 |
Querying Shortest Paths on Time Dependent Road Networks |
2019 |
VLDB |
17 |
9.7886535e-05 |
| 1,805 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
37 |
9.5938877e-05 |
| 1,993 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
27 |
9.2348951e-05 |
| 2,188 |
Graph Stream Summarization: From Big Bang to Big Crunch |
2016 |
SIGMOD |
16 |
8.8874856e-05 |
| 2,349 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
22 |
8.5969105e-05 |
| 2,424 |
BigDansing: A System for Big Data Cleansing |
2015 |
SIGMOD |
34 |
8.483813e-05 |
| 2,467 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
19 |
8.4171827e-05 |
| 2,514 |
Deep Learning for Blocking in Entity Matching: A Design Space Exploration |
2021 |
VLDB |
25 |
8.3610263e-05 |
| 2,518 |
Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation |
2022 |
VLDB |
37 |
8.3532841e-05 |
| 2,784 |
Interaction between Record Matching and Data Repairing |
2011 |
SIGMOD |
17 |
8.0147123e-05 |
| 2,967 |
Towards Dependable Data Repairing with Fixing Rules |
2014 |
SIGMOD |
18 |
7.8016942e-05 |
| 3,296 |
Lightning Fast and Space Efficient Inequality Joins |
2015 |
VLDB |
14 |
7.444648e-05 |
| 3,402 |
RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - |
2018 |
VLDB |
22 |
7.3269775e-05 |
| 3,421 |
UGuide – User-Guided Discovery of FD-Detectable Errors |
2017 |
SIGMOD |
13 |
7.3120636e-05 |
| 3,473 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
15 |
7.2697311e-05 |
| 3,489 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
21 |
7.2594757e-05 |
| 3,634 |
NADEEF/ER: Generic and Interactive Entity Resolution |
2014 |
SIGMOD |
7 |
7.1454896e-05 |
| 3,684 |
Generating Concise Entity Matching Rules |
2017 |
SIGMOD |
9 |
7.0983735e-05 |
| 3,942 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
13 |
6.9105312e-05 |
| 3,967 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
8 |
6.887577e-05 |
| 4,113 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
12 |
6.7964307e-05 |
| 4,411 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
10 |
6.6089506e-05 |
| 4,690 |
Learned Cardinality Estimation for Similarity Queries |
2021 |
SIGMOD |
18 |
6.4667478e-05 |
| 4,845 |
Pattern Functional Dependencies for Data Cleaning |
2020 |
VLDB |
13 |
6.3808938e-05 |
| 4,865 |
Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks |
2021 |
SIGMOD |
14 |
6.3740341e-05 |
| 4,869 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
18 |
6.3719187e-05 |
| 4,955 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
9 |
6.3386868e-05 |
| 5,255 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
12 |
6.2057681e-05 |
| 5,391 |
ANMAT: Automatic Knowledge Discovery and Error Detection through Pattern Functional Dependencies |
2019 |
SIGMOD |
5 |
6.1495611e-05 |
| 5,499 |
HAIChart: Human and AI Paired Visualization System |
2024 |
VLDB |
7 |
6.1027393e-05 |
| 5,515 |
KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing |
2015 |
VLDB |
9 |
6.0962265e-05 |
| 5,645 |
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks |
2023 |
VLDB |
7 |
6.0509471e-05 |
| 5,703 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
7 |
6.02892e-05 |
| 5,760 |
DeepEye: Creating Good Data Visualizations by Keyword Search |
2018 |
SIGMOD |
9 |
6.0014434e-05 |
| 5,815 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
13 |
5.9824897e-05 |
| 5,851 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
10 |
5.9692535e-05 |
| 6,154 |
NADEEF: A Generalized Data Cleaning System |
2013 |
VLDB |
13 |
5.8665355e-05 |
| 6,159 |
Automatic Data Acquisition for Deep Learning |
2021 |
VLDB |
7 |
5.8653048e-05 |
| 6,684 |
Controllable Tabular Data Synthesis Using Diffusion Models |
2024 |
SIGMOD |
4 |
5.70774e-05 |
| 6,878 |
Towards Democratizing Relational Data Visualization |
2019 |
SIGMOD |
4 |
5.6563449e-05 |
| 7,092 |
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning |
2026 |
VLDB |
2 |
5.5991152e-05 |
| 7,203 |
Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning |
2023 |
VLDB |
8 |
5.5845207e-05 |
| 7,262 |
Learned Data-aware Image Representations of Line Charts for Similarity Search |
2023 |
SIGMOD |
7 |
5.5694763e-05 |
| 7,273 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
9 |
5.5668569e-05 |
| 7,293 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
7 |
5.5616488e-05 |
| 7,589 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
4 |
5.4880216e-05 |
| 7,618 |
Are Large Language Models a Good Replacement of Taxonomies? |
2024 |
VLDB |
3 |
5.481538e-05 |
| 7,654 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5 |
5.4746904e-05 |
| 7,771 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
2 |
5.453953e-05 |
| 7,854 |
CoClean: Collaborative Data Cleaning |
2020 |
SIGMOD |
4 |
5.4380042e-05 |
| 8,483 |
DADER: Hands-Off Entity Resolution with Domain Adaptation |
2022 |
VLDB |
3 |
5.3311145e-05 |
| 9,187 |
CerFix: A System for Cleaning Data with Certain Fixes |
2011 |
VLDB |
3 |
5.2099641e-05 |
| 9,220 |
Natural Language to SQL: State of the Art and Open Problems |
2025 |
VLDB |
6 |
5.2036865e-05 |
| 9,320 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
6 |
5.1941278e-05 |
| 9,481 |
VisClean: Interactive Cleaning for Progressive Visualization |
2020 |
VLDB |
8 |
5.1689938e-05 |
| 9,626 |
Interactive and Deterministic Data Cleaning: A Tossed Stone Raises a Thousand Ripples |
2016 |
SIGMOD |
1 |
5.1477881e-05 |
| 9,682 |
Debugging Large-Scale Data Science Pipelines using Dagger |
2020 |
VLDB |
3 |
5.1413762e-05 |
| 10,165 |
Rheem: Enabling Multi-Platform Task Execution |
2016 |
SIGMOD |
7 |
5.0664415e-05 |
| 10,298 |
Andromeda: Debugging Database Performance Issues with Retrieval-Augmented Large Language Models |
2025 |
SIGMOD |
1 |
5.0407989e-05 |
| 10,814 |
Document-to-Database: Extraction Meets Relational Semantics |
2026 |
VLDB |
1 |
4.9769913e-05 |
| 11,278 |
Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation |
2025 |
VLDB |
3 |
4.9769913e-05 |
| 12,086 |
Interactively Discovering and Ranking Desired Tuples without Writing SQL Queries |
2020 |
SIGMOD |
3 |
4.9769913e-05 |
| 13,597 |
DataMosaic: An Interactive Demonstration of Constraint-Driven Document-to-Database Construction |
2026 |
VLDB |
0 |
- |
| 13,804 |
DeepTrack: Monitoring and Exploring Spatio-Temporal Data - A Case of Tracking COVID-19 - |
2020 |
VLDB |
3 |
- |
| 13,859 |
Errata for “Lightning Fast and Space Efficient Inequality Joins” (PVLDB 8(13): 2074-2085) |
2017 |
VLDB |
0 |
- |