| 489 |
Distributed Representations of Tuples for Entity Resolution |
2018 |
VLDB |
0.0001761456 |
| 725 |
NADEEF: A Commodity Data Cleaning System |
2013 |
SIGMOD |
0.00014617251 |
| 963 |
The Data Civilizer System |
2017 |
CIDR |
0.00012935145 |
| 998 |
Towards Certain Fixes with Editing Rules and Master Data |
2010 |
VLDB |
0.0001275238 |
| 1,101 |
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing |
2015 |
SIGMOD |
0.00012168934 |
| 1,128 |
Graph Pattern Matching: From Intractable to Polynomial Time |
2010 |
VLDB |
0.0001206219 |
| 1,351 |
Detecting Data Errors: Where are we and what needs to be done? |
2016 |
VLDB |
0.00011064851 |
| 1,465 |
Synthesizing Entity Matching Rules by Examples |
2018 |
VLDB |
0.00010689571 |
| 1,550 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00010385904 |
| 1,745 |
Querying Shortest Paths on Time Dependent Road Networks |
2019 |
VLDB |
9.861709e-05 |
| 2,019 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
9.2983994e-05 |
| 2,148 |
Graph Stream Summarization: From Big Bang to Big Crunch |
2016 |
SIGMOD |
9.0825156e-05 |
| 2,297 |
Raha: A Configuration-Free Error Detection System |
2019 |
SIGMOD |
8.7897221e-05 |
| 2,398 |
BigDansing: A System for Big Data Cleansing |
2015 |
SIGMOD |
8.631172e-05 |
| 2,475 |
Deep Learning for Blocking in Entity Matching: A Design Space Exploration |
2021 |
VLDB |
8.5277654e-05 |
| 2,723 |
Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation |
2022 |
VLDB |
8.2049453e-05 |
| 2,748 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
8.1707811e-05 |
| 2,782 |
Interaction between Record Matching and Data Repairing |
2011 |
SIGMOD |
8.1280264e-05 |
| 2,852 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
8.0455088e-05 |
| 2,981 |
Towards Dependable Data Repairing with Fixing Rules |
2014 |
SIGMOD |
7.8960114e-05 |
| 3,295 |
Lightning Fast and Space Efficient Inequality Joins |
2015 |
VLDB |
7.5477715e-05 |
| 3,351 |
RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - |
2018 |
VLDB |
7.4937347e-05 |
| 3,380 |
UGuide – User-Guided Discovery of FD-Detectable Errors |
2017 |
SIGMOD |
7.457539e-05 |
| 3,436 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
7.4157897e-05 |
| 3,611 |
NADEEF/ER: Generic and Interactive Entity Resolution |
2014 |
SIGMOD |
7.2580469e-05 |
| 3,713 |
Generating Concise Entity Matching Rules |
2017 |
SIGMOD |
7.176496e-05 |
| 3,787 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
7.1249098e-05 |
| 3,886 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
7.0460597e-05 |
| 4,617 |
Learned Cardinality Estimation for Similarity Queries |
2021 |
SIGMOD |
6.604437e-05 |
| 4,757 |
Pattern Functional Dependencies for Data Cleaning |
2020 |
VLDB |
6.5220005e-05 |
| 4,808 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
6.4991553e-05 |
| 4,809 |
Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks |
2021 |
SIGMOD |
6.4984309e-05 |
| 4,835 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
6.486592e-05 |
| 5,268 |
ANMAT: Automatic Knowledge Discovery and Error Detection through Pattern Functional Dependencies |
2019 |
SIGMOD |
6.2923403e-05 |
| 5,366 |
HAIChart: Human and AI Paired Visualization System |
2024 |
VLDB |
6.2457679e-05 |
| 5,397 |
KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing |
2015 |
VLDB |
6.232136e-05 |
| 5,404 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
6.2294192e-05 |
| 5,518 |
Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks |
2023 |
VLDB |
6.1885722e-05 |
| 5,634 |
DeepEye: Creating Good Data Visualizations by Keyword Search |
2018 |
SIGMOD |
6.1412942e-05 |
| 5,644 |
A Demo of the Data Civilizer System |
2017 |
SIGMOD |
6.1378898e-05 |
| 5,682 |
RetClean: Retrieval-Based Data Cleaning Using LLMs and Data Lakes |
2024 |
VLDB |
6.123486e-05 |
| 5,730 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
6.1071585e-05 |
| 5,856 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
6.0644656e-05 |
| 6,050 |
NADEEF: A Generalized Data Cleaning System |
2013 |
VLDB |
5.9945626e-05 |
| 6,074 |
Automatic Data Acquisition for Deep Learning |
2021 |
VLDB |
5.9860133e-05 |
| 6,555 |
Controllable Tabular Data Synthesis Using Diffusion Models |
2024 |
SIGMOD |
5.8415111e-05 |
| 6,737 |
Towards Democratizing Relational Data Visualization |
2019 |
SIGMOD |
5.7875328e-05 |
| 7,112 |
Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning |
2023 |
VLDB |
5.6990782e-05 |
| 7,199 |
Learned Data-aware Image Representations of Line Charts for Similarity Search |
2023 |
SIGMOD |
5.6759927e-05 |
| 7,483 |
Are Large Language Models a Good Replacement of Taxonomies? |
2024 |
VLDB |
5.6074551e-05 |