Database Paper Browser

Back to papers

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Summary: Hybrid KV+hidden-state cache expands batch size under GPU memory limits for LLM inference. Adaptive scheduling with formal optimization and guarantees tunes batch composition, yielding up to 8.8× throughput vs SOTA on 13B–66B models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
7283
Venue
SIGMOD
Year
2025
Pagerank
4.3006524e-05
Overall Rank
9,677 | 32.75%
DOI
10.1145/3725394

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,222 RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference 2026 VLDB 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 35 of 35 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
332 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00027173479
1,162 Sancus: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks 2022 VLDB 0.00013573136
1,500 Analyzing and Mitigating Data Stalls in DNN Training 2021 VLDB 0.00011636174
2,165 Accelerating Large Scale Real-Time GNN Inference using Channel Pruning 2021 VLDB 9.3925908e-05
2,425 DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU 2023 SIGMOD 8.8414587e-05
2,678 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.3224016e-05
3,291 Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics 2021 VLDB 7.2607192e-05
3,466 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.0645718e-05
3,694 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 6.8316905e-05
3,715 Zebra: When Temporal Graph Neural Networks Meet Temporal Personalized PageRank 2023 VLDB 6.8176818e-05
4,054 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 6.4910804e-05
4,175 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 6.3772575e-05
5,062 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 5.7172262e-05
5,169 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 5.642415e-05
5,210 Self-Tuning Query Scheduling for Analytical Workloads 2021 SIGMOD 5.6244961e-05
5,332 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 5.5640779e-05
5,485 ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic Graphs 2024 VLDB 5.4817019e-05
6,350 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 5.0957601e-05
6,361 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 5.0903244e-05
6,479 EARLY: Efficient and Reliable Graph Neural Network for Dynamic Graphs 2023 SIGMOD 5.0405101e-05
7,014 SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement 2024 SIGMOD 4.8570865e-05
7,143 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 4.8143774e-05
7,286 DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning 2024 VLDB 4.7701372e-05
7,378 Distribution-Based Query Scheduling 2013 VLDB 4.7428007e-05
7,559 Transaction Scheduling: From Conflicts to Runtime Conflicts 2023 SIGMOD 4.7065517e-05
7,687 Towards Optimal Transaction Scheduling 2024 VLDB 4.6745165e-05
8,057 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 4.5903427e-05
8,117 SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 4.5788485e-05
9,374 MorphStream: Adaptive Scheduling for Scalable Transactional Stream Processing on Multicores 2023 SIGMOD 4.3453721e-05
9,704 ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program Reuse 2022 VLDB 4.2953234e-05
9,785 Capsule*: An Out-of-Core Training Mechanism for Colossal GNNs 2025 SIGMOD 4.2799988e-05
9,786 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 4.2799988e-05
9,787 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 4.2799988e-05
9,788 Demonstration of Accelerating Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 4.2799988e-05
13,164 STile: Searching Hybrid Sparse Formats for Sparse Deep Learning Operators Automatically 2024 SIGMOD -
Previous Page 1 / 1 Next

Semantically Similar Papers