AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft
Summary: AutoToken predicts peak resource usage for recurring big-data queries in serverless analytics. A lightweight, scalable predictor using multiple query-plan identifiers to detect recurring templates, integrated with Peregrine and validated on SCOPE jobs. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Rathijit Sen (Microsoft)
- 2. Alekh Jindal (Microsoft)
- 3. Hiren Patel (Microsoft)
- 4. Shi Qiao (Microsoft)
BibTeX Citation
@article{sen_vldb20,
title = {{AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft}},
author = {Sen, Rathijit and Jindal, Alekh and Patel, Hiren and Qiao, Shi},
journal = {PVLDB},
series = {{VLDB} '20},
volume = {13},
number = {12},
pages = {3326--3339},
doi = {10.14778/3415478.3415554},
url = {https://doi.org/10.14778/3415478.3415554},
year = {2020}
}
Incoming Citations (Sorted by Pagerank)
Showing 10 of 10 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 30 | SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets | 2008 | VLDB | 0.00051174276 |
| 895 | Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance | 2010 | VLDB | 0.00013357681 |
| 923 | Starfish: A Self-tuning System for Big Data Analytics | 2011 | CIDR | 0.00013189886 |
| 1,468 | Towards a Learning Optimizer for Shared Clouds | 2019 | VLDB | 0.00010686496 |
| 2,822 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings | 2020 | SIGMOD | 8.0898536e-05 |
| 3,605 | Computation Reuse in Analytics Job Service at Microsoft | 2018 | SIGMOD | 7.2640711e-05 |
| 9,854 | SparkCruise: Handsfree Computation Reuse in Spark | 2019 | VLDB | 5.2091816e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,809 | Auto-WLM: Machine Learning Enhanced Workload Management in Amazon Redshift | 2023 | SIGMOD |
| 2 | 5,865 | Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems | 2019 | VLDB |
| 3 | 5,676 | Scalable Progressive Analytics on Big Data in the Cloud | 2013 | VLDB |
| 4 | 11,222 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service | 2024 | VLDB |
| 5 | 3,953 | Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload | 2021 | SIGMOD |
| 6 | 2,822 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings | 2020 | SIGMOD |
| 7 | 6,946 | KEA: Tuning an Exabyte-Scale Data Infrastructure | 2021 | SIGMOD |
| 8 | 4,742 | Continuous Cloud-Scale Query Optimization and Processing | 2013 | VLDB |
| 9 | 5,059 | Steering Query Optimizers: A Practical Take on Big Data Workloads | 2021 | SIGMOD |
| 10 | 6,102 | AutoExecutor: Predictive Parallelism for Spark SQL Queries | 2021 | VLDB |