| 2,348 |
The Dawn of Natural Language to SQL: Are We Fully Ready? |
2024 |
VLDB |
8.6009821e-05 |
| 2,395 |
A Learned Query Rewrite System using Monte Carlo Tree Search |
2022 |
VLDB |
8.5281914e-05 |
| 2,484 |
AnalyticDB: Real-time OLAP Database System at Alibaba Cloud |
2019 |
VLDB |
8.3993882e-05 |
| 2,690 |
Cost-based or Learning-based? A Hybrid Query Optimizer for Query Plan Selection |
2022 |
VLDB |
8.1258173e-05 |
| 3,741 |
FACE: A Normalizing Flow based Cardinality Estimator |
2022 |
VLDB |
7.0594076e-05 |
| 3,941 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
6.9138042e-05 |
| 4,409 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
6.6120807e-05 |
| 4,482 |
LearnedSQLGen: Constraint-aware SQL Generation using Reinforcement Learning |
2022 |
SIGMOD |
6.5802486e-05 |
| 4,623 |
Cost-Effective Crowdsourced Entity Resolution: A Partial-Order Approach |
2016 |
SIGMOD |
6.4978811e-05 |
| 4,731 |
QUEST: Query Optimization in Unstructured Document Analysis |
2025 |
VLDB |
6.4462032e-05 |
| 4,864 |
Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL Benchmarks |
2021 |
SIGMOD |
6.3770469e-05 |
| 4,867 |
Selective Data Acquisition in the Wild for Model Charging |
2022 |
VLDB |
6.3749346e-05 |
| 5,061 |
CDB: A Crowd-Powered Database System |
2018 |
VLDB |
6.2912543e-05 |
| 5,814 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
5.9852808e-05 |
| 5,848 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
5.9720806e-05 |
| 6,156 |
Automatic Data Acquisition for Deep Learning |
2021 |
VLDB |
5.8680826e-05 |
| 6,869 |
Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents |
2025 |
VLDB |
5.659905e-05 |
| 7,042 |
Human-in-the-loop Outlier Detection |
2020 |
SIGMOD |
5.6142998e-05 |
| 7,090 |
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning |
2026 |
VLDB |
5.601767e-05 |
| 7,201 |
Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning |
2023 |
VLDB |
5.5871656e-05 |
| 7,258 |
Learned Data-aware Image Representations of Line Charts for Similarity Search |
2023 |
SIGMOD |
5.572114e-05 |
| 7,583 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
5.4906208e-05 |
| 7,648 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5.4772833e-05 |
| 7,762 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
5.456536e-05 |
| 8,476 |
DADER: Hands-Off Entity Resolution with Domain Adaptation |
2022 |
VLDB |
5.333639e-05 |
| 8,632 |
DocDB: A Database for Unstructured Document Analysis |
2025 |
VLDB |
5.2985603e-05 |
| 8,680 |
Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark |
2026 |
VLDB |
5.2905577e-05 |
| 8,800 |
PACE: Poisoning Attacks on Learned Cardinality Estimation |
2024 |
SIGMOD |
5.2742531e-05 |
| 9,210 |
Natural Language to SQL: State of the Art and Open Problems |
2025 |
VLDB |
5.2061511e-05 |
| 9,470 |
VisClean: Interactive Cleaning for Progressive Visualization |
2020 |
VLDB |
5.1714419e-05 |
| 9,583 |
CDB: Optimizing Queries with Crowd-Based Selections and Joins |
2017 |
SIGMOD |
5.1571823e-05 |
| 10,711 |
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs |
2026 |
VLDB |
4.9793485e-05 |
| 10,821 |
Data-efficient Online Training for Direct Alignment in LLMs |
2026 |
VLDB |
4.9793485e-05 |
| 10,837 |
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries |
2026 |
VLDB |
4.9793485e-05 |
| 11,213 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
4.9793485e-05 |
| 12,080 |
Interactively Discovering and Ranking Desired Tuples without Writing SQL Queries |
2020 |
SIGMOD |
4.9793485e-05 |