QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark
Summary: Query-merging via microRDD converts many small Spark queries into a few larger ones; queries embedded as data enable shared inputs. Dynamic partition sizing minimizes runtime overhead, yielding 10.6x-36.6x speedups over Spark for small queries. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yeonsu Park (Pohang University of Science and Technology)
- 2. Byungchul Tak (Kyungpook National University)
- 3. Wook-Shin Han (Pohang University of Science and Technology)
BibTeX Citation
@inproceedings{park_sigmod23,
title = {{QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark}},
author = {Park, Yeonsu and Tak, Byungchul and Han, Wook-Shin},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589279},
url = {https://dl.acm.org/doi/10.1145/3589279},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 2 | 12,141 | A Demonstration of AQWA: Adaptive Query-Workload-Aware Partitioning of Big Spatial Data | 2015 | VLDB |
| 3 | 8,175 | SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft | 2021 | VLDB |
| 4 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |
| 5 | 9,655 | [Demo] Low-latency Spark Queries on Updatable Data | 2019 | SIGMOD |
| 6 | 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 7 | 7,777 | S2RDF: RDF Querying with SPARQL on Spark | 2016 | VLDB |
| 8 | 8,439 | New Query Optimization Techniques in the Spark Engine of Azure Synapse | 2022 | VLDB |
| 9 | 8,615 | A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning | 2024 | VLDB |
| 10 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |