Sieve: A Learned Data-Skipping Index for Data Analytics
Summary: Sieve is a learned data-skipping index that models block-distribution trends over the key space with piecewise-linear functions. It groups keys with similar distributions to trade index size for false positives, reducing accessed blocks by up to 80% and query time by 42%. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yulai Tong (Huazhong University of Science and Technology)
- 2. Jiazhen Liu (Huazhong University of Science and Technology)
- 3. Hua Wang (Huazhong University of Science and Technology)
- 4. Ke Zhou (Huazhong University of Science and Technology)
- 5. Rongfeng He (Huawei)
- 6. Qin Zhang (Huawei)
- 7. Cheng Wang (Huawei)
BibTeX Citation
@article{tong_vldb23,
title = {{Sieve: A Learned Data-Skipping Index for Data Analytics}},
author = {Tong, Yulai and Liu, Jiazhen and Wang, Hua and Zhou, Ke and He, Rongfeng and Zhang, Qin and Wang, Cheng},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {11},
pages = {3214--3226},
doi = {10.14778/3611479.3611520},
url = {https://doi.org/10.14778/3611479.3611520},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,368 | Accelerating String-key Learned Index Structures via Memoization-based Incremental Training | 2024 | VLDB | 5.6315363e-05 |
| 7,469 | Optimizing Collections of Bloom Filters within a Space Budget | 2024 | VLDB | 5.609743e-05 |
| 10,378 | High Performance or Low Memory? An Updatable Learned Index Framework for Time-Space Tradeoff | 2026 | SIGMOD | 5.093636e-05 |
| 10,529 | Robust Predicate Transfer with Dynamic Execution | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,532 | Stacked Filters: Learning to Filter by Structure | 2021 | VLDB |
| 2 | 7,368 | Accelerating String-key Learned Index Structures via Memoization-based Incremental Training | 2024 | VLDB |
| 3 | 6,653 | Adaptive Data Skipping in Main-Memory Systems | 2016 | SIGMOD |
| 4 | 4,782 | Cuckoo Index: A Lightweight Secondary Index Structure | 2020 | VLDB |
| 5 | 10,672 | Optimizing Block Skipping for High-Dimensional Data with Learned Adaptive Curve | 2025 | SIGMOD |
| 6 | 12,191 | A Partitioning Framework for Aggressive Data Skipping | 2014 | VLDB |
| 7 | 9,498 | S3: A Scalable In-memory Skip-List Index for Key-Value Store | 2019 | VLDB |
| 8 | 8,424 | Sieve: A Middleware Approach to Scalable Access Control for Database Management Systems | 2020 | VLDB |
| 9 | 9,145 | SIEVE: Effective Filtered Vector Search with Collection of Indexes | 2025 | VLDB |
| 10 | 10,329 | SieveSketch: A Fine-grained and Adaptive Sketch Framework for Accurate Frequency Estimation | 2026 | SIGMOD |