| 756 |
CodeS: Towards Building Open-source Language Models for Text-to-SQL |
2024 |
SIGMOD |
0.0001431656 |
| 1,550 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00010385904 |
| 2,019 |
RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data Preparation |
2021 |
VLDB |
9.2983994e-05 |
| 2,260 |
Online Topic-Aware Influence Maximization |
2015 |
VLDB |
8.8526838e-05 |
| 2,748 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
8.1707811e-05 |
| 3,029 |
Trajectory Simplification: An Experimental Study and Quality Analysis |
2018 |
VLDB |
7.8340364e-05 |
| 3,262 |
iCrowd: An Adaptive Crowdsourcing Framework |
2015 |
SIGMOD |
7.5846052e-05 |
| 3,436 |
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration |
2023 |
SIGMOD |
7.4157897e-05 |
| 3,787 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
7.1249098e-05 |
| 3,886 |
GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete Data |
2023 |
SIGMOD |
7.0460597e-05 |
| 4,261 |
Relational Data Synthesis using Generative Adversarial Networks: A Design Space Exploration |
2020 |
VLDB |
6.7982037e-05 |
| 4,835 |
Adaptive Data Augmentation for Supervised Learning over Missing Data |
2021 |
VLDB |
6.486592e-05 |
| 4,984 |
SEAL: Spatio-Textual Similarity Search |
2012 |
VLDB |
6.412239e-05 |
| 5,192 |
Discovering Your Selling Points: Personalized Social Influential Tags Exploration |
2017 |
SIGMOD |
6.3263349e-05 |
| 5,211 |
CDB: A Crowd-Powered Database System |
2018 |
VLDB |
6.3161892e-05 |
| 5,730 |
Domain Adaptation for Deep Entity Resolution |
2022 |
SIGMOD |
6.1071585e-05 |
| 5,856 |
Automatic Database Configuration Debugging using Retrieval-Augmented Language Models |
2025 |
SIGMOD |
6.0644656e-05 |
| 6,115 |
Fine-grained Concept Linking using Neural Networks in Healthcare |
2018 |
SIGMOD |
5.9709282e-05 |
| 6,128 |
GEMINI: An Integrative Healthcare Analytics System |
2014 |
VLDB |
5.9668307e-05 |
| 6,555 |
Controllable Tabular Data Synthesis Using Diffusion Models |
2024 |
SIGMOD |
5.8415111e-05 |
| 6,862 |
DBease: Making Databases User-friendly and Easily Accessible |
2011 |
CIDR |
5.7514733e-05 |
| 6,908 |
Cost-Effective Data Annotation using Game-Based Crowdsourcing |
2019 |
VLDB |
5.7415834e-05 |
| 7,357 |
Crowdsourced Data Management: Overview and Challenges |
2017 |
SIGMOD |
5.6346837e-05 |
| 8,177 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
5.4730821e-05 |
| 8,336 |
DADER: Hands-Off Entity Resolution with Domain Adaptation |
2022 |
VLDB |
5.451516e-05 |
| 8,495 |
CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling |
2019 |
SIGMOD |
5.4139128e-05 |
| 8,797 |
Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table Representations |
2024 |
SIGMOD |
5.3702298e-05 |
| 9,177 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
5.3078984e-05 |
| 9,401 |
CDB: Optimizing Queries with Crowd-Based Selections and Joins |
2017 |
SIGMOD |
5.2755515e-05 |
| 9,550 |
Data Imputation with Limited Data Redundancy Using Data Lakes |
2025 |
VLDB |
5.2528121e-05 |
| 9,777 |
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking |
2025 |
VLDB |
5.2209769e-05 |
| 9,870 |
Natural Language to SQL: State of the Art and Open Problems |
2025 |
VLDB |
5.2043672e-05 |
| 10,285 |
Reward-SQL: Boosting Text-to-SQL via Stepwise Execution-Aware Reasoning and Process-Supervised Rewards |
2026 |
SIGMOD |
5.093636e-05 |
| 10,304 |
VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis] |
2026 |
SIGMOD |
5.093636e-05 |
| 10,537 |
TACO: A Benchmark for Open-Domain Text-to-SQL with Ambiguous and Cross-Database Queries |
2026 |
VLDB |
5.093636e-05 |
| 10,705 |
Andromeda: Debugging Database Performance Issues with Retrieval-Augmented Large Language Models |
2025 |
SIGMOD |
5.093636e-05 |
| 10,867 |
Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation |
2025 |
VLDB |
5.093636e-05 |
| 10,931 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
5.093636e-05 |
| 11,211 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5.093636e-05 |
| 11,236 |
Improving Graph Compression for Efficient Resource-Constrained Graph Analytics |
2024 |
VLDB |
5.093636e-05 |
| 11,545 |
OpenTFV: An Open Domain Table-Based Fact Verification System |
2022 |
SIGMOD |
5.093636e-05 |
| 12,236 |
TsingNUS: A Location-Based Service System Towards Live City |
2013 |
SIGMOD |
5.093636e-05 |