Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGs
Summary: Introduces Pasta, a cost-based optimizer for scheduling pipelined execution of dataflow DAGs. General across cost models, it exploits cost-function structure to reduce materialization and improve plan quality on real-world workflows. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Xiaozhen Liu (University of California Irvine)
- 2. Yicong Huang (University of California Irvine)
- 3. Xinyuan Lin (University of California Irvine)
- 4. Avinash Kumar (University of California Irvine)
- 5. Sadeem Alsudais (King Saud University)
- 6. Chen Li (University of California Irvine)
BibTeX Citation
@inproceedings{liu_sigmod24,
title = {{Pasta: A Cost-Based Optimizer for Generating Pipelining Schedules for Dataflow DAGs}},
author = {Liu, Xiaozhen and Huang, Yicong and Lin, Xinyuan and Kumar, Avinash and Alsudais, Sadeem and Li, Chen},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3698832},
url = {https://dl.acm.org/doi/10.1145/3698832},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,845 | Explaining Outputs in Modern Data Analytics | 2016 | VLDB |
| 2 | 2,131 | Optimization Algorithms for Exploiting the Parallelism-Communication Tradeoff in Pipelined Parallelism | 1994 | VLDB |
| 3 | 2,164 | Opening the Black Boxes in Data Flow Optimization | 2012 | VLDB |
| 4 | 2,657 | Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities | 2021 | SIGMOD |
| 5 | 8,237 | Meta-Dataflows: Efficient Exploratory Dataflow Jobs | 2018 | SIGMOD |
| 6 | 3,275 | Optimizing Analytic Data Flows for Multiple Execution Engines | 2012 | SIGMOD |
| 7 | 822 | Pipelining in Multi-Query Optimization | 2001 | PODS |
| 8 | 6,309 | Materialization and Reuse Optimizations for Production Data Science Pipelines | 2022 | SIGMOD |
| 9 | 2,196 | Spinning Fast Iterative Data Flows | 2012 | VLDB |
| 10 | 4,117 | Schedule Optimization for Data Processing Flows on the Cloud | 2011 | SIGMOD |