DBScholar

Back to papers

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Summary: Hybrid KV+hidden-state cache expands batch size under GPU memory limits for LLM inference. Adaptive scheduling with formal optimization and guarantees tunes batch composition, yielding up to 8.8× throughput vs SOTA on 13B–66B models. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h01d34dd363184bc4
Venue
SIGMOD
Year
2025
Pagerank
5.8631131e-05
Overall Rank
6,166 | 58.55%
DOI
10.1145/3725394

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{gao_sigmod25,
        title = {{Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving}},
        author = {Gao, Shihong and Zhang, Xin and Shen, Yanyan and Chen, Lei},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725394},
        url = {https://dl.acm.org/doi/10.1145/3725394},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 35 of 35 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
281 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022295232
1,134 SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks 2022 VLDB 0.00011893521
1,475 Analyzing and Mitigating Data Stalls in DNN Training 2021 VLDB 0.00010556672
2,439 Accelerating Large Scale Real-Time GNN Inference using Channel Pruning 2021 VLDB 8.4652234e-05
2,506 DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU 2023 SIGMOD 8.3774747e-05
2,519 HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework 2022 VLDB 8.3521095e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8742664e-05
3,017 Zebra: When Temporal Graph Neural Networks Meet Temporal Personalized PageRank 2023 VLDB 7.7471344e-05
3,219 Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics 2021 VLDB 7.5207996e-05
3,399 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 7.3349076e-05
3,601 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 7.1757877e-05
3,811 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 7.0087478e-05
3,969 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8899861e-05
4,351 Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity 2024 VLDB 6.642803e-05
4,388 Self-Tuning Query Scheduling for Analytical Workloads 2021 SIGMOD 6.6228033e-05
4,498 PQCache: Product Quantization-based KVCache for Long Context LLM Inference 2025 SIGMOD 6.5762771e-05
4,764 ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic Graphs 2024 VLDB 6.4276208e-05
4,825 DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning 2024 VLDB 6.3923554e-05
4,945 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.3435891e-05
4,963 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.3384091e-05
5,009 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3158866e-05
6,324 EARLY: Efficient and Reliable Graph Neural Network for Dynamic Graphs 2023 SIGMOD 5.8137796e-05
6,826 SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement 2024 SIGMOD 5.6700316e-05
6,872 Distribution-Based Query Scheduling 2013 VLDB 5.6592113e-05
7,003 Towards Optimal Transaction Scheduling 2024 VLDB 5.6244707e-05
7,280 Transaction Scheduling: From Conflicts to Runtime Conflicts 2023 SIGMOD 5.5671848e-05
7,978 Biathlon: Harnessing Model Resilience for Accelerating ML Inference Pipelines 2024 VLDB 5.4136835e-05
8,251 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3687311e-05
9,530 MorphStream: Adaptive Scheduling for Scalable Transactional Stream Processing on Multicores 2023 SIGMOD 5.1653263e-05
10,013 ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program Reuse 2022 VLDB 5.0971861e-05
10,041 Capsule*: An Out-of-Core Training Mechanism for Colossal GNNs 2025 SIGMOD 5.0921006e-05
10,042 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0921006e-05
10,043 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0921006e-05
10,044 Demonstration of Accelerating Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 5.0921006e-05
13,671 STile: Searching Hybrid Sparse Formats for Sparse Deep Learning Operators Automatically 2024 SIGMOD -
Previous Page 1 / 1 Next

Semantically Similar Papers