The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap –
Summary: Matryoshka enables nested parallelism in dataflow engines with a two-phase flattening that turns programs into flat ones, even with inner control flow. It adds nesting primitives and runtime data-aware optimizations, validated on PageRank and K-means. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Gábor E. Gévay (Technical University of Berlin)
- 2. Jorge-Arnulfo Quiané-Ruiz (German National Research Center for Information Technology; Technical University of Berlin)
- 3. Volker Markl (German National Research Center for Information Technology; Technical University of Berlin)
BibTeX Citation
@inproceedings{gevay_sigmod21,
title = {{The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap –}},
author = {Gévay, Gábor E. and Quiané-Ruiz, Jorge-Arnulfo and Markl, Volker},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457287},
url = {https://dl.acm.org/doi/10.1145/3448016.3457287},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,301 | Babelfish: Efficient Execution of Polyglot Queries | 2022 | VLDB | 6.2750553e-05 |
| 6,538 | UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads | 2022 | VLDB | 5.8477764e-05 |
| 7,160 | DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines | 2022 | CIDR | 5.6855887e-05 |
| 7,232 | Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications | 2023 | SIGMOD | 5.6659017e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 12 of 12 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,196 | Massively Parallel Data Analysis with PACTs on Nephele | 2010 | VLDB |
| 2 | 954 | Parallel Evaluation of Conjunctive Queries | 2011 | PODS |
| 3 | 10,771 | Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins | 2025 | SIGMOD |
| 4 | 8,738 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB |
| 5 | 8,237 | Meta-Dataflows: Efficient Exploratory Dataflow Jobs | 2018 | SIGMOD |
| 6 | 12,237 | Iterative Parallel Data Processing with Stratosphere: An Inside Look | 2013 | SIGMOD |
| 7 | 2,717 | Implicit Parallelism through Deep Language Embedding | 2015 | SIGMOD |
| 8 | 6,377 | Scalable Querying of Nested Data | 2021 | VLDB |
| 9 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |
| 10 | 2,196 | Spinning Fast Iterative Data Flows | 2012 | VLDB |