| 1,245 |
Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing |
2025 |
CIDR |
0.00011507415 |
| 1,550 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
0.00010385904 |
| 1,700 |
Aria: A Fast and Practical Deterministic OLTP Database |
2020 |
VLDB |
9.9726588e-05 |
| 2,748 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
8.1707811e-05 |
| 2,888 |
AI Meets Database: AI4DB and DB4AI |
2021 |
SIGMOD |
7.9941489e-05 |
| 3,031 |
Interactive Outlier Exploration in Big Data Streams |
2014 |
VLDB |
7.8324917e-05 |
| 3,787 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
7.1249098e-05 |
| 3,869 |
Smile: A System to Support Machine Learning on EEG Data at Scale |
2019 |
VLDB |
7.0609879e-05 |
| 4,062 |
AutoOD: Automatic Outlier Detection |
2023 |
SIGMOD |
6.9309994e-05 |
| 4,582 |
A Demonstration of AutoOD: A Self-Tuning Anomaly Detection System |
2022 |
VLDB |
6.6193306e-05 |
| 4,869 |
Epoch-based Commit and Replication in Distributed OLTP Databases |
2021 |
VLDB |
6.4709867e-05 |
| 5,228 |
Pluto: Sample Selection for Robust Anomaly Detection on Polluted Log Data |
2024 |
SIGMOD |
6.3082722e-05 |
| 5,340 |
Machine Learning for Databases |
2021 |
VLDB |
6.2603359e-05 |
| 5,404 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
6.2294192e-05 |
| 5,846 |
Continuously Adaptive Similarity Search |
2020 |
SIGMOD |
6.0679671e-05 |
| 6,371 |
QUEST: Query Optimization in Unstructured Document Analysis |
2025 |
VLDB |
5.8962187e-05 |
| 6,494 |
Extract-Transform-Load for Video Streams |
2023 |
VLDB |
5.8639535e-05 |
| 6,766 |
Sharing-Aware Outlier Analytics over High-Volume Data Streams |
2016 |
SIGMOD |
5.7802844e-05 |
| 6,899 |
Human-in-the-loop Outlier Detection |
2020 |
SIGMOD |
5.7431609e-05 |
| 7,613 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
5.5817761e-05 |
| 7,893 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
5.5209866e-05 |
| 8,170 |
Efficient Discovery of Sequence Outlier Patterns |
2019 |
VLDB |
5.4741508e-05 |
| 8,242 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
5.459288e-05 |
| 8,340 |
Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents |
2025 |
VLDB |
5.4487212e-05 |
| 8,802 |
Outlier Summarization via Human Interpretable Rules |
2024 |
VLDB |
5.3685765e-05 |
| 8,876 |
LANCET: Labeling Complex Data at Scale |
2021 |
VLDB |
5.3534114e-05 |
| 8,996 |
High Performance Stream Query Processing With Correlation-Aware Partitioning |
2014 |
VLDB |
5.3357805e-05 |
| 9,177 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
5.3078984e-05 |
| 9,543 |
OIE: An Interpretable System for Outlier Explanation and Summarization |
2025 |
SIGMOD |
5.2528121e-05 |
| 9,631 |
Lingua Manga: A Generic Large Language Model Centric System for Data Curation |
2023 |
VLDB |
5.2434488e-05 |
| 9,764 |
Complex Event Analytics: Online Aggregation of Stream Sequence Patterns |
2014 |
SIGMOD |
5.2233526e-05 |
| 10,034 |
HARMONY: A Scalable Distributed Vector Database for High-Throughput Approximate Nearest Neighbor Search |
2026 |
SIGMOD |
5.173224e-05 |
| 10,255 |
HotHash: Hotness-Aware Consistent Hashing for Cloud Databases |
2026 |
SIGMOD |
5.093636e-05 |
| 10,497 |
Scalable Clustering Over High Dimensional Vector Streams |
2026 |
SIGMOD |
5.093636e-05 |
| 10,527 |
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs |
2026 |
VLDB |
5.093636e-05 |
| 10,623 |
KEN: An Execution Engine for Unstructured Database Systems |
2026 |
VLDB |
5.093636e-05 |
| 10,658 |
Agree to Disagree: Robust Anomaly Detection with Noisy Labels |
2025 |
SIGMOD |
5.093636e-05 |
| 10,719 |
Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents |
2025 |
SIGMOD |
5.093636e-05 |
| 10,800 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
5.093636e-05 |
| 10,931 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
5.093636e-05 |
| 11,168 |
RITA: Group Attention is All You Need for Timeseries Analytics |
2024 |
SIGMOD |
5.093636e-05 |
| 11,211 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5.093636e-05 |
| 11,219 |
MetaStore: Analyzing Deep Learning Meta-Data at Scale |
2024 |
VLDB |
5.093636e-05 |
| 11,712 |
ATLANTIC: Making Database Differentially Private and Faster with Accuracy Guarantee |
2021 |
VLDB |
5.093636e-05 |
| 11,887 |
SWIFT: Mining Representative Patterns from Large Event Streams |
2019 |
VLDB |
5.093636e-05 |
| 13,339 |
DocDB: A Database for Unstructured Document Analysis |
2025 |
VLDB |
- |