| 748 |
Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing |
2025 |
CIDR |
0.00014281926 |
| 1,392 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00010807936 |
| 1,693 |
Aria: A Fast and Practical Deterministic OLTP Database |
2020 |
VLDB |
9.8567076e-05 |
| 2,467 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
8.4211526e-05 |
| 2,908 |
AI Meets Database: AI4DB and DB4AI |
2021 |
SIGMOD |
7.8742664e-05 |
| 3,043 |
Interactive Outlier Exploration in Big Data Streams |
2014 |
VLDB |
7.7212998e-05 |
| 3,489 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
7.2627807e-05 |
| 3,955 |
Smile: A System to Support Machine Learning on EEG Data at Scale |
2019 |
VLDB |
6.9025583e-05 |
| 4,103 |
AutoOD: Automatic Outlier Detection |
2023 |
SIGMOD |
6.8066074e-05 |
| 4,409 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
6.6120807e-05 |
| 4,686 |
A Demonstration of AutoOD: A Self-Tuning Anomaly Detection System |
2022 |
VLDB |
6.4708106e-05 |
| 4,731 |
QUEST: Query Optimization in Unstructured Document Analysis |
2025 |
VLDB |
6.4462032e-05 |
| 4,741 |
Machine Learning for Databases |
2021 |
VLDB |
6.4410027e-05 |
| 4,820 |
Epoch-based Commit and Replication in Distributed OLTP Databases |
2021 |
VLDB |
6.3962005e-05 |
| 5,252 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
6.2087063e-05 |
| 5,351 |
Pluto: Sample Selection for Robust Anomaly Detection on Polluted Log Data |
2024 |
SIGMOD |
6.1667315e-05 |
| 5,935 |
Continuously Adaptive Similarity Search |
2020 |
SIGMOD |
5.9419167e-05 |
| 6,583 |
Extract-Transform-Load for Video Streams |
2023 |
VLDB |
5.7438181e-05 |
| 6,869 |
Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents |
2025 |
VLDB |
5.659905e-05 |
| 6,904 |
Sharing-Aware Outlier Analytics over High-Volume Data Streams |
2016 |
SIGMOD |
5.6506055e-05 |
| 7,042 |
Human-in-the-loop Outlier Detection |
2020 |
SIGMOD |
5.6142998e-05 |
| 7,270 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
5.5694935e-05 |
| 7,291 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
5.5642829e-05 |
| 7,648 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5.4772833e-05 |
| 7,762 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
5.456536e-05 |
| 7,924 |
Outlier Summarization via Human Interpretable Rules |
2024 |
VLDB |
5.425954e-05 |
| 8,344 |
Efficient Discovery of Sequence Outlier Patterns |
2019 |
VLDB |
5.3513288e-05 |
| 8,632 |
DocDB: A Database for Unstructured Document Analysis |
2025 |
VLDB |
5.2985603e-05 |
| 8,680 |
Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark |
2026 |
VLDB |
5.2905577e-05 |
| 8,848 |
HARMONY: A Scalable Distributed Vector Database for High-Throughput Approximate Nearest Neighbor Search |
2026 |
SIGMOD |
5.2646236e-05 |
| 9,036 |
LANCET: Labeling Complex Data at Scale |
2021 |
VLDB |
5.2332952e-05 |
| 9,154 |
High Performance Stream Query Processing With Correlation-Aware Partitioning |
2014 |
VLDB |
5.2176811e-05 |
| 9,311 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
5.1965878e-05 |
| 9,725 |
OIE: An Interpretable System for Outlier Explanation and Summarization |
2025 |
SIGMOD |
5.1349531e-05 |
| 9,808 |
Lingua Manga: A Generic Large Language Model Centric System for Data Curation |
2023 |
VLDB |
5.1257999e-05 |
| 9,942 |
Complex Event Analytics: Online Aggregation of Stream Sequence Patterns |
2014 |
SIGMOD |
5.1061569e-05 |
| 10,468 |
HotHash: Hotness-Aware Consistent Hashing for Cloud Databases |
2026 |
SIGMOD |
4.9793485e-05 |
| 10,684 |
Scalable Clustering Over High Dimensional Vector Streams |
2026 |
SIGMOD |
4.9793485e-05 |
| 10,711 |
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs |
2026 |
VLDB |
4.9793485e-05 |
| 10,821 |
Data-efficient Online Training for Direct Alignment in LLMs |
2026 |
VLDB |
4.9793485e-05 |
| 10,837 |
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries |
2026 |
VLDB |
4.9793485e-05 |
| 11,069 |
KEN: An Execution Engine for Unstructured Database Systems |
2026 |
VLDB |
4.9793485e-05 |
| 11,101 |
Agree to Disagree: Robust Anomaly Detection with Noisy Labels |
2025 |
SIGMOD |
4.9793485e-05 |
| 11,151 |
Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents |
2025 |
SIGMOD |
4.9793485e-05 |
| 11,213 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
4.9793485e-05 |
| 11,513 |
RITA: Group Attention is All You Need for Timeseries Analytics |
2024 |
SIGMOD |
4.9793485e-05 |
| 11,556 |
MetaStore: Analyzing Deep Learning Meta-Data at Scale |
2024 |
VLDB |
4.9793485e-05 |
| 12,016 |
ATLANTIC: Making Database Differentially Private and Faster with Accuracy Guarantee |
2021 |
VLDB |
4.9793485e-05 |
| 12,187 |
SWIFT: Mining Representative Patterns from Large Event Streams |
2019 |
VLDB |
4.9793485e-05 |