| 533 |
CodeS: Towards Building Open-source Language Models for Text-to-SQL |
2024 |
SIGMOD |
41 |
0.00016818605 |
| 1,387 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
19 |
0.00010829741 |
| 1,993 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
27 |
9.2348951e-05 |
| 2,309 |
Online Topic-Aware Influence Maximization |
2015 |
VLDB |
15 |
8.6639186e-05 |
| 2,467 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
19 |
8.4171827e-05 |
| 3,010 |
Trajectory Simplification: An Experimental Study and Quality Analysis |
2018 |
VLDB |
8 |
7.7541653e-05 |
| 3,199 |
iCrowd: An Adaptive Crowdsourcing Framework |
2015 |
SIGMOD |
23 |
7.5419863e-05 |
| 3,473 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
15 |
7.2697311e-05 |
| 3,489 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
21 |
7.2594757e-05 |
| 3,942 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
13 |
6.9105312e-05 |
| 4,113 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
12 |
6.7964307e-05 |
| 4,329 |
Relational Data Synthesis using Generative Adversarial Networks: A Design Space Exploration |
2020 |
VLDB |
12 |
6.656734e-05 |
| 4,955 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
9 |
6.3386868e-05 |
| 5,050 |
CDB: A Crowd-Powered Database System |
2018 |
VLDB |
10 |
6.2955765e-05 |
| 5,115 |
SEAL: Spatio-Textual Similarity Search |
2012 |
VLDB |
8 |
6.2654746e-05 |
| 5,304 |
Discovering Your Selling Points: Personalized Social Influential Tags Exploration |
2017 |
SIGMOD |
6 |
6.1867379e-05 |
| 5,815 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
13 |
5.9824897e-05 |
| 5,851 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
10 |
5.9692535e-05 |
| 6,236 |
Fine-grained Concept Linking using Neural Networks in Healthcare |
2018 |
SIGMOD |
6 |
5.83935e-05 |
| 6,263 |
GEMINI: An Integrative Healthcare Analytics System |
2014 |
VLDB |
5 |
5.8301898e-05 |
| 6,684 |
Controllable Tabular Data Synthesis Using Diffusion Models |
2024 |
SIGMOD |
4 |
5.70774e-05 |
| 6,983 |
DBease: Making Databases User-friendly and Easily Accessible |
2011 |
CIDR |
2 |
5.6261545e-05 |
| 7,054 |
Cost-Effective Data Annotation using Game-Based Crowdsourcing |
2019 |
VLDB |
8 |
5.6101007e-05 |
| 7,273 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
9 |
5.5668569e-05 |
| 7,505 |
Crowdsourced Data Management: Overview and Challenges |
2017 |
SIGMOD |
3 |
5.5057737e-05 |
| 7,589 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
4 |
5.4880216e-05 |
| 7,654 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5 |
5.4746904e-05 |
| 7,819 |
Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations |
2024 |
SIGMOD |
4 |
5.4460313e-05 |
| 8,483 |
DADER: Hands-Off Entity Resolution with Domain Adaptation |
2022 |
VLDB |
3 |
5.3311145e-05 |
| 8,643 |
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking |
2025 |
VLDB |
2 |
5.2956539e-05 |
| 8,672 |
CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling |
2019 |
SIGMOD |
2 |
5.2901608e-05 |
| 9,220 |
Natural Language to SQL: State of the Art and Open Problems |
2025 |
VLDB |
6 |
5.2036865e-05 |
| 9,320 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
6 |
5.1941278e-05 |
| 9,591 |
CDB: Optimizing Queries with Crowd-Based Selections and Joins |
2017 |
SIGMOD |
9 |
5.154741e-05 |
| 10,279 |
Improving Graph Compression for Efficient Resource-Constrained Graph Analytics |
2024 |
VLDB |
6 |
5.0461162e-05 |
| 10,298 |
Andromeda: Debugging Database Performance Issues with Retrieval-Augmented Large Language Models |
2025 |
SIGMOD |
1 |
5.0407989e-05 |
| 10,508 |
Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards |
2026 |
SIGMOD |
7 |
4.9769913e-05 |
| 10,527 |
VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis] |
2026 |
SIGMOD |
0 |
4.9769913e-05 |
| 10,729 |
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries |
2026 |
VLDB |
1 |
4.9769913e-05 |
| 10,878 |
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation |
2026 |
VLDB |
1 |
4.9769913e-05 |
| 11,278 |
Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation |
2025 |
VLDB |
3 |
4.9769913e-05 |
| 11,860 |
OpenTFV: An Open Domain Table-Based Fact Verification System |
2022 |
SIGMOD |
3 |
4.9769913e-05 |
| 12,533 |
TsingNUS: A Location-Based Service System Towards Live City |
2013 |
SIGMOD |
1 |
4.9769913e-05 |
| 13,607 |
DA-Studio: An Agentic System for End-to-End Data Analysis |
2026 |
VLDB |
0 |
- |