| 457 |
Distributed Representations of Tuples for Entity Resolution |
2018 |
VLDB |
0.00017907103 |
| 697 |
NADEEF: A Commodity Data Cleaning System |
2013 |
SIGMOD |
0.00014694048 |
| 972 |
The Data Civilizer System |
2017 |
CIDR |
0.00012763234 |
| 989 |
Towards Certain Fixes with Editing Rules and Master Data |
2010 |
VLDB |
0.0001265344 |
| 1,099 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
0.00012037058 |
| 1,148 |
Graph Pattern Matching: From Intractable to Polynomial Time |
2010 |
VLDB |
0.00011809728 |
| 1,344 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00010956518 |
| 1,392 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00010807936 |
| 1,493 |
Synthesizing Entity Matching Rules by Examples |
2018 |
VLDB |
0.00010505174 |
| 1,722 |
Querying Shortest Paths on Time Dependent Road Networks |
2019 |
VLDB |
9.7932895e-05 |
| 1,805 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
9.59842e-05 |
| 1,991 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
9.2383849e-05 |
| 2,186 |
Graph Stream Summarization: From Big Bang to Big Crunch |
2016 |
SIGMOD |
8.8916948e-05 |
| 2,348 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
8.6009821e-05 |
| 2,423 |
BigDansing: A System for Big Data Cleansing |
2015 |
SIGMOD |
8.4877894e-05 |
| 2,467 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
8.4211526e-05 |
| 2,514 |
Deep Learning for Blocking in Entity Matching: A Design Space Exploration |
2021 |
VLDB |
8.3648432e-05 |
| 2,522 |
Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation |
2022 |
VLDB |
8.3477168e-05 |
| 2,784 |
Interaction between Record Matching and Data Repairing |
2011 |
SIGMOD |
8.018482e-05 |
| 2,964 |
Towards Dependable Data Repairing with Fixing Rules |
2014 |
SIGMOD |
7.8052551e-05 |
| 3,295 |
Lightning Fast and Space Efficient Inequality Joins |
2015 |
VLDB |
7.448168e-05 |
| 3,402 |
RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - |
2018 |
VLDB |
7.3304477e-05 |
| 3,422 |
UGuide – User-Guided Discovery of FD-Detectable Errors |
2017 |
SIGMOD |
7.3138834e-05 |
| 3,473 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
7.2728706e-05 |
| 3,489 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
7.2627807e-05 |
| 3,634 |
NADEEF/ER: Generic and Interactive Entity Resolution |
2014 |
SIGMOD |
7.1482079e-05 |
| 3,681 |
Generating Concise Entity Matching Rules |
2017 |
SIGMOD |
7.1014441e-05 |
| 3,941 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
6.9138042e-05 |
| 3,973 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
6.8876964e-05 |
| 4,114 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
6.7971182e-05 |
| 4,409 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
6.6120807e-05 |
| 4,688 |
Learned Cardinality Estimation for Similarity Queries |
2021 |
SIGMOD |
6.4697463e-05 |
| 4,843 |
Pattern Functional Dependencies for Data Cleaning |
2020 |
VLDB |
6.3839159e-05 |
| 4,864 |
Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks |
2021 |
SIGMOD |
6.3770469e-05 |
| 4,867 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
6.3749346e-05 |
| 4,953 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
6.3416889e-05 |
| 5,252 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
6.2087063e-05 |
| 5,384 |
ANMAT: Automatic Knowledge Discovery and Error Detection through Pattern Functional Dependencies |
2019 |
SIGMOD |
6.1524649e-05 |
| 5,495 |
HAIChart: Human and AI Paired Visualization System |
2024 |
VLDB |
6.1056297e-05 |
| 5,512 |
KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing |
2015 |
VLDB |
6.0991137e-05 |
| 5,643 |
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks |
2023 |
VLDB |
6.0538129e-05 |
| 5,700 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
6.0317753e-05 |
| 5,759 |
DeepEye: Creating Good Data Visualizations by Keyword Search |
2018 |
SIGMOD |
6.0042797e-05 |
| 5,814 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
5.9852808e-05 |
| 5,848 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
5.9720806e-05 |
| 6,152 |
NADEEF: A Generalized Data Cleaning System |
2013 |
VLDB |
5.8693135e-05 |
| 6,156 |
Automatic Data Acquisition for Deep Learning |
2021 |
VLDB |
5.8680826e-05 |
| 6,680 |
Controllable Tabular Data Synthesis Using Diffusion Models |
2024 |
SIGMOD |
5.7104433e-05 |
| 6,875 |
Towards Democratizing Relational Data Visualization |
2019 |
SIGMOD |
5.6589849e-05 |
| 7,090 |
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning |
2026 |
VLDB |
5.601767e-05 |