DBScholar

Back to papers

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Summary: Hybrid KV+hidden-state cache expands batch size under GPU memory limits for LLM inference. Adaptive scheduling with formal optimization and guarantees tunes batch composition, yielding up to 8.8× throughput vs SOTA on 13B–66B models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h01d34dd363184bc4
Venue
SIGMOD
Year
2025
Pagerank
5.8603375e-05
Overall Rank
6,168 | 58.55%
DOI
10.1145/3725394

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{gao_sigmod25,
        title = {{Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving}},
        author = {Gao, Shihong and Zhang, Xin and Shen, Yanyan and Chen, Lei},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725394},
        url = {https://dl.acm.org/doi/10.1145/3725394},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 35 of 35 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
282 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022302793
1,134 SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks 2022 VLDB 0.0001188789
1,475 Analyzing and Mitigating Data Stalls in DNN Training 2021 VLDB 0.00010551698
2,440 Accelerating Large Scale Real-Time GNN Inference using Channel Pruning 2021 VLDB 8.461216e-05
2,506 DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU 2023 SIGMOD 8.3735089e-05
2,521 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.3481558e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8716173e-05
3,018 Zebra: When Temporal Graph Neural Networks Meet Temporal Personalized PageRank 2023 VLDB 7.743467e-05
3,220 Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics 2021 VLDB 7.5172795e-05
3,399 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 7.3314353e-05
3,601 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 7.1724084e-05
3,812 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 7.0056415e-05
3,971 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8868815e-05
4,352 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.6396584e-05
4,391 Self-Tuning Query Scheduling for Analytical Workloads 2021 SIGMOD 6.6196806e-05
4,500 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 6.573164e-05
4,767 ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic Graphs 2024 VLDB 6.4245781e-05
4,827 DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning 2024 VLDB 6.3893293e-05
4,949 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.3405861e-05
4,965 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.3354086e-05
5,011 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3131083e-05
6,328 EARLY: Efficient and Reliable Graph Neural Network for Dynamic Graphs 2023 SIGMOD 5.8110274e-05
6,831 SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement 2024 SIGMOD 5.6673474e-05
6,876 Distribution-Based Query Scheduling 2013 VLDB 5.6565323e-05
7,005 Towards Optimal Transaction Scheduling 2024 VLDB 5.6218081e-05
7,283 Transaction Scheduling: From Conflicts to Runtime Conflicts 2023 SIGMOD 5.5645493e-05
7,983 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.4111208e-05
8,257 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3661896e-05
9,539 MorphStream: Adaptive Scheduling for Scalable Transactional Stream Processing on Multicores 2023 SIGMOD 5.1628811e-05
10,018 ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program Reuse 2022 VLDB 5.0947732e-05
10,046 Capsule*: An Out-of-Core Training Mechanism for Colossal GNNs 2025 SIGMOD 5.0896901e-05
10,047 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0896901e-05
10,048 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0896901e-05
10,049 Demonstration of Accelerating Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 5.0896901e-05
13,676 STile: Searching Hybrid Sparse Formats for Sparse Deep Learning Operators Automatically 2024 SIGMOD -
Previous Page 1 / 1 Next

Semantically Similar Papers