DBScholar

Back to papers

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Summary: Hybrid KV+hidden-state cache expands batch size under GPU memory limits for LLM inference. Adaptive scheduling with formal optimization and guarantees tunes batch composition, yielding up to 8.8× throughput vs SOTA on 13B–66B models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
7344
Venue
SIGMOD
Year
2025
Pagerank
5.7430032e-05
Overall Rank
6,900 | 52.67%
DOI
10.1145/3725394

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{gao_sigmod25,
        title = {{Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving}},
        author = {Gao, Shihong and Zhang, Xin and Shen, Yanyan and Chen, Lei},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725394},
        url = {https://dl.acm.org/doi/10.1145/3725394},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 35 of 35 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
295 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022238183
1,132 SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks 2022 VLDB 0.00012041292
1,446 Analyzing and Mitigating Data Stalls in DNN Training 2021 VLDB 0.0001076818
2,386 Accelerating Large Scale Real-Time GNN Inference using Channel Pruning 2021 VLDB 8.6490185e-05
2,485 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.5145736e-05
2,668 DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU 2023 SIGMOD 8.2750247e-05
2,888 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.9941489e-05
3,210 Zebra: When Temporal Graph Neural Networks Meet Temporal Personalized PageRank 2023 VLDB 7.6352864e-05
3,253 Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics 2021 VLDB 7.5936939e-05
3,541 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 7.3280673e-05
3,634 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 7.2358691e-05
3,764 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 7.1446931e-05
4,114 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8941194e-05
4,430 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 6.7091071e-05
4,493 Self-Tuning Query Scheduling for Analytical Workloads 2021 SIGMOD 6.6647555e-05
4,912 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.4481656e-05
4,956 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.4290135e-05
5,003 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.4065691e-05
5,205 ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic Graphs 2024 VLDB 6.318626e-05
6,261 EARLY: Efficient and Reliable Graph Neural Network for Dynamic Graphs 2023 SIGMOD 5.9366952e-05
6,730 SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement 2024 SIGMOD 5.7895038e-05
7,075 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 5.7099047e-05
7,087 DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning 2024 VLDB 5.7069166e-05
7,149 Distribution-Based Query Scheduling 2013 VLDB 5.6878283e-05
7,159 Transaction Scheduling: From Conflicts to Runtime Conflicts 2023 SIGMOD 5.6857508e-05
7,334 Towards Optimal Transaction Scheduling 2024 VLDB 5.6425501e-05
7,847 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.5330423e-05
8,101 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.4870581e-05
9,366 MorphStream: Adaptive Scheduling for Scalable Transactional Stream Processing on Multicores 2023 SIGMOD 5.2789847e-05
9,826 ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program Reuse 2022 VLDB 5.2141422e-05
9,879 Capsule*: An Out-of-Core Training Mechanism for Colossal GNNs 2025 SIGMOD 5.2040783e-05
9,880 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.2040783e-05
9,881 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.2040783e-05
9,882 Demonstration of Accelerating Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 5.2040783e-05
13,354 STile: Searching Hybrid Sparse Formats for Sparse Deep Learning Operators Automatically 2024 SIGMOD -
Previous Page 1 / 1 Next

Semantically Similar Papers