| 717 |
Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing |
2025 |
CIDR |
46 |
0.00014528323 |
| 1,387 |
Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes |
2023 |
CIDR |
19 |
0.00010829741 |
| 1,694 |
Aria: A Fast and Practical Deterministic OLTP Database |
2020 |
VLDB |
44 |
9.8520416e-05 |
| 2,467 |
Few-shot Text-to-SQL Translation using Structure and Content Prompt Learning |
2023 |
SIGMOD |
19 |
8.4171827e-05 |
| 2,908 |
AI Meets Database: AI4DB and DB4AI |
2021 |
SIGMOD |
31 |
7.8716173e-05 |
| 3,044 |
Interactive Outlier Exploration in Big Data Streams |
2014 |
VLDB |
8 |
7.7177446e-05 |
| 3,489 |
Combining Small Language Models and Large Language Models for Zero-Shot NL2SQL |
2024 |
VLDB |
21 |
7.2594757e-05 |
| 3,956 |
Smile: A System to Support Machine Learning on EEG Data at Scale |
2019 |
VLDB |
6 |
6.8992907e-05 |
| 4,105 |
AutoOD: Automatic Outlier Detection |
2023 |
SIGMOD |
8 |
6.8033852e-05 |
| 4,411 |
LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes |
2024 |
VLDB |
10 |
6.6089506e-05 |
| 4,565 |
QUEST: Query Optimization in Unstructured Document Analysis |
2025 |
VLDB |
9 |
6.5295131e-05 |
| 4,689 |
A Demonstration of AutoOD: A Self-Tuning Anomaly Detection System |
2022 |
VLDB |
6 |
6.4677473e-05 |
| 4,743 |
Machine Learning for Databases |
2021 |
VLDB |
15 |
6.4379536e-05 |
| 4,822 |
Epoch-based Commit and Replication in Distributed OLTP Databases |
2021 |
VLDB |
13 |
6.3931726e-05 |
| 5,255 |
Dagger: A Data (not code) Debugger |
2020 |
CIDR |
12 |
6.2057681e-05 |
| 5,357 |
Pluto: Sample Selection for Robust Anomaly Detection on Polluted Log Data |
2024 |
SIGMOD |
3 |
6.1638123e-05 |
| 5,935 |
Continuously Adaptive Similarity Search |
2020 |
SIGMOD |
5 |
5.9391038e-05 |
| 6,586 |
Extract-Transform-Load for Video Streams |
2023 |
VLDB |
9 |
5.741099e-05 |
| 6,874 |
Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents |
2025 |
VLDB |
5 |
5.6572257e-05 |
| 6,906 |
Sharing-Aware Outlier Analytics over High-Volume Data Streams |
2016 |
SIGMOD |
4 |
5.6479305e-05 |
| 7,043 |
Human-in-the-loop Outlier Detection |
2020 |
SIGMOD |
17 |
5.611642e-05 |
| 7,273 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
9 |
5.5668569e-05 |
| 7,293 |
Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics |
2019 |
VLDB |
7 |
5.5616488e-05 |
| 7,654 |
MisDetect: Iterative Mislabel Detection using Early Loss |
2024 |
VLDB |
5 |
5.4746904e-05 |
| 7,771 |
LakeCompass: An End-to-End System for Data Maintenance, Search and Analysis in Data Lakes |
2024 |
VLDB |
2 |
5.453953e-05 |
| 7,928 |
Outlier Summarization via Human Interpretable Rules |
2024 |
VLDB |
4 |
5.4233854e-05 |
| 8,182 |
DocDB: A Database for Unstructured Document Analysis |
2025 |
VLDB |
3 |
5.3810756e-05 |
| 8,347 |
Efficient Discovery of Sequence Outlier Patterns |
2019 |
VLDB |
2 |
5.3487955e-05 |
| 8,688 |
Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark |
2026 |
VLDB |
1 |
5.2880532e-05 |
| 8,858 |
HARMONY: A Scalable Distributed Vector Database for High-Throughput Approximate Nearest Neighbor Search |
2026 |
SIGMOD |
2 |
5.2621314e-05 |
| 9,044 |
LANCET: Labeling Complex Data at Scale |
2021 |
VLDB |
3 |
5.2308178e-05 |
| 9,163 |
High Performance Stream Query Processing With Correlation-Aware Partitioning |
2014 |
VLDB |
2 |
5.2152138e-05 |
| 9,320 |
VerifAI: Verified Generative AI |
2024 |
CIDR |
6 |
5.1941278e-05 |
| 9,730 |
OIE: An Interpretable System for Outlier Explanation and Summarization |
2025 |
SIGMOD |
1 |
5.1325223e-05 |
| 9,815 |
Lingua Manga: A Generic Large Language Model Centric System for Data Curation |
2023 |
VLDB |
1 |
5.1233734e-05 |
| 9,949 |
Complex Event Analytics: Online Aggregation of Stream Sequence Patterns |
2014 |
SIGMOD |
7 |
5.1037398e-05 |
| 10,479 |
HotHash: Hotness-Aware Consistent Hashing for Cloud Databases |
2026 |
SIGMOD |
0 |
4.9769913e-05 |
| 10,695 |
Scalable Clustering Over High Dimensional Vector Streams |
2026 |
SIGMOD |
0 |
4.9769913e-05 |
| 10,721 |
BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 10,831 |
Data-efficient Online Training for Direct Alignment in LLMs |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 10,847 |
EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 11,078 |
KEN: An Execution Engine for Unstructured Database Systems |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 11,110 |
Agree to Disagree: Robust Anomaly Detection with Noisy Labels |
2025 |
SIGMOD |
0 |
4.9769913e-05 |
| 11,160 |
Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents |
2025 |
SIGMOD |
0 |
4.9769913e-05 |
| 11,222 |
Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset Selection |
2025 |
SIGMOD |
1 |
4.9769913e-05 |
| 11,519 |
RITA: Group Attention is All You Need for Timeseries Analytics |
2024 |
SIGMOD |
0 |
4.9769913e-05 |
| 11,562 |
MetaStore: Analyzing Deep Learning Meta-Data at Scale |
2024 |
VLDB |
0 |
4.9769913e-05 |
| 12,022 |
ATLANTIC: Making Database Differentially Private and Faster with Accuracy Guarantee |
2021 |
VLDB |
0 |
4.9769913e-05 |
| 12,193 |
SWIFT: Mining Representative Patterns from Large Event Streams |
2019 |
VLDB |
0 |
4.9769913e-05 |