DBScholar

Back to papers

Runtime Variation in Big Data Analytics

Summary: Two-step predictor for runtime distribution: shape features plus a classifier with >96% accuracy. First large-scale study predicting enterprise analytics runtime categories; enables what-if analyses on allocation and scheduling. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h3321c7a84daf9095
Venue
SIGMOD
Year
2023
Pagerank
5.4256585e-05
Overall Rank
7,925 | 46.72%
DOI
10.1145/3588921

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{zhu_sigmod23,
        title = {{Runtime Variation in Big Data Analytics}},
        author = {Zhu, Yiwen and Sen, Rathijit and Horton, Robert and Agosta, John Mark},
        series = {{SIGMOD} '23},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3588921},
        url = {https://dl.acm.org/doi/10.1145/3588921},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
7,982 InTime: Towards Performance Predictability In Byzantine Fault Tolerant Proof-of-Stake Consensus 2025 SIGMOD 5.4131553e-05
10,207 From Logs to Causal Inference: Diagnosing Large Systems 2025 VLDB 5.0596605e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
78 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00036684414
314 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning 2019 SIGMOD 0.00021282642
892 Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance 2010 VLDB 0.00013227162
1,918 Predictable Performance for Unpredictable Workloads 2009 VLDB 9.3789552e-05
2,446 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4547121e-05
3,193 A Statistical Perspective on Discovering Functional Dependencies in Noisy Data 2020 SIGMOD 7.5505481e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2986853e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9393157e-05
4,470 Recurring Job Optimization in Scope 2012 SIGMOD 6.5853561e-05
4,818 Continuous Cloud-Scale Query Optimization and Processing 2013 VLDB 6.3964573e-05
4,912 A Top-Down Approach to Achieving Performance Predictability in Database Systems 2017 SIGMOD 6.3578361e-05
6,196 AutoExecutor: Predictive Parallelism for Spark SQL Queries 2021 VLDB 5.8533869e-05
6,245 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.837329e-05
6,958 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 5.634024e-05
7,084 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 5.6029455e-05
7,759 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 5.4575614e-05
9,358 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.1869771e-05
Previous Page 1 / 1 Next

Semantically Similar Papers