PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
Summary: PystachIO repurposes PyTorch tensor runtimes for distributed, storage-resident OLAP on GPU clusters with RDMA and NVMe. Its execution pipeline overlaps computation, network, and storage I/O, achieving up to 3× speedups over distributed GPU query engines. (summarized by gpt-5.6-luna on Aug 17 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Jigao Luo (Technical University of Darmstadt)
- 2. Nils Boeschen (German National Research Center for Information Technology; Technical University of Darmstadt)
- 3. Muhammad El-Hindi (Technical University of Munich)
- 4. Carsten Binnig (German National Research Center for Information Technology; Hessian Center for Artificial Intelligence; Technical University of Darmstadt)
BibTeX Citation
@article{luo_vldb26,
title = {{PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks \& Fast Storage}},
author = {Luo, Jigao and Boeschen, Nils and El-Hindi, Muhammad and Binnig, Carsten},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {9},
pages = {2494--2507},
doi = {10.14778/3819518.3819566},
url = {https://doi.org/10.14778/3819518.3819566},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,020 | Share the Tensor Tea: How Databases can Leverage the Machine Learning Ecosystem | 2022 | VLDB |
| 2 | 5,541 | Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics | 2025 | VLDB |
| 3 | 2,246 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel | 2023 | VLDB |
| 4 | 5,371 | Distributed GPU Joins on Fast RDMA-capable Networks | 2023 | SIGMOD |
| 5 | 4,645 | Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs | 2021 | VLDB |
| 6 | 6,516 | GPU-accelerated data management under the test of time | 2020 | CIDR |
| 7 | 9,863 | GPU Acceleration of SQL Analytics on Compressed Data | 2026 | VLDB |
| 8 | 522 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB |
| 9 | 5,991 | Terabyte-Scale Analytics in the Blink of an Eye | 2026 | VLDB |
| 10 | 2,460 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB |