DBScholar

Back to papers

Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects

Summary: NVLink 2.0-based interconnect eliminates CPU-GPU transfer bottlenecks, enabling large-scale in-memory processing on GPUs. Demonstrates scalable no-partitioning hash join beyond GPU memory with up to 18x speedup vs PCIe 3.0 and 7.3x vs optimized CPU. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5982
Venue
SIGMOD
Year
2020
Pagerank
8.4116562e-05
Overall Rank
2,566 | 82.40%
DOI
10.1145/3318464.3389705

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{lutz_sigmod20,
        title = {{Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects}},
        author = {Lutz, Clemens and Breß, Sebastian and Zeuch, Steffen and Rabl, Tilmann and Markl, Volker},
        series = {{SIGMOD} '20},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3318464.3389705},
        url = {https://dl.acm.org/doi/10.1145/3318464.3389705},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 30 of 30 citing papers.

Rank Citing Paper Year Venue Pagerank
2,823 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.0893814e-05
3,890 GaccO - A GPU-accelerated OLTP DBMS 2022 SIGMOD 7.0435206e-05
4,089 TCUDB: Accelerating Database with Tensor Processors 2022 SIGMOD 6.9096857e-05
4,144 Tile-based Lightweight Integer Compression in GPU 2022 SIGMOD 6.8744592e-05
4,243 Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS 2022 VLDB 6.8073248e-05
4,535 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 6.6419266e-05
4,861 CXL and the Return of Scale-Up Database Engines 2024 VLDB 6.4747307e-05
5,422 GPU Database Systems Characterization and Optimization 2024 VLDB 6.2245373e-05
5,587 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1596139e-05
5,807 Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs 2021 VLDB 6.0805143e-05
6,065 Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics 2025 VLDB 5.9895288e-05
6,774 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7778738e-05
7,058 BOSS - An Architecture for Database Kernel Composition 2024 VLDB 5.7154673e-05
7,160 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.6855887e-05
7,235 An Examination of CXL Memory Use Cases for In-Memory Database Management Systems using SAP HANA 2024 VLDB 5.6655351e-05
7,432 Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs 2025 VLDB 5.6199783e-05
7,591 Efficiently Processing Joins and Grouped Aggregations on GPUs 2025 SIGMOD 5.5900315e-05
7,760 NOCAP: Near-Optimal Correlation-Aware Partitioning Joins 2023 SIGMOD 5.5505651e-05
7,820 Parallelizing Intra-Window Join on Multicores: An Experimental Study 2021 SIGMOD 5.5373345e-05
7,959 Terabyte-Scale Analytics in the Blink of an Eye 2026 VLDB 5.5181056e-05
8,443 Analyzing Vectorized Hash Tables Across CPU Architectures 2023 VLDB 5.4243766e-05
8,596 Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs 2024 SIGMOD 5.4051515e-05
8,808 Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs 2024 VLDB 5.3661351e-05
9,333 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.2887551e-05
9,376 H-Rocks: CPU-GPU accelerated Heterogeneous RocksDB on Persistent Memory 2025 SIGMOD 5.2755515e-05
9,873 Workload Placement on Heterogeneous CPU-GPU Systems 2024 VLDB 5.2043672e-05
10,260 L3: A GPU-Native Co-Designed Data Format for Learned Lossless Lightweight Compression 2026 SIGMOD 5.093636e-05
10,409 TQEx: Tensor-based Query Engine Enhanced by Bridging the Gap 2026 SIGMOD 5.093636e-05
10,580 GPU Acceleration of SQL Analytics on Compressed Data 2026 VLDB 5.093636e-05
11,231 Accelerating Merkle Patricia Trie with GPU 2024 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 31 of 31 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
209 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024932174
241 Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age 2014 SIGMOD 0.00023654664
278 FAST: Fast Architecture Sensitive Tree Search on Modern CPUs and GPUs 2010 SIGMOD 0.00022476841
360 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020182846
631 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015591241
823 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00013792901
892 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013376761
912 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013269804
1,203 NUMA-aware algorithms: the case of data shuffling 2013 CIDR 0.00011671169
1,265 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011415709
1,452 Fast Computation of Database Operations using Graphics Processors 2004 SIGMOD 0.00010745803
1,574 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010321274
1,644 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010132912
1,666 HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics 2016 VLDB 0.00010068964
1,977 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines 2019 VLDB 9.3641101e-05
2,140 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0991487e-05
2,232 Database Compression on Graphics Processors 2010 VLDB 8.8970926e-05
2,667 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.2756346e-05
2,713 Robust Query Processing in Co-Processor-accelerated Databases 2016 SIGMOD 8.2122726e-05
3,060 SABER: Window-Based Hybrid Stream Processing for Heterogeneous Architectures 2016 SIGMOD 7.8062155e-05
3,378 A Hybrid B+-tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms 2016 SIGMOD 7.4587887e-05
3,471 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.3885861e-05
3,682 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.2035518e-05
3,782 CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers 2019 VLDB 7.1298979e-05
3,791 Hardware-conscious Query Processing in GPU-accelerated Analytical Engines 2019 CIDR 7.1235328e-05
3,895 The Case For Heterogeneous HTAP 2017 CIDR 7.0416054e-05
4,376 Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources 2017 VLDB 6.7364442e-05
6,143 ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation 2019 VLDB 5.9622615e-05
6,572 A Morsel-Driven Query Execution Engine for Heterogeneous Multi-Cores 2019 VLDB 5.8373744e-05
6,835 GPU-accelerated data management under the test of time 2020 CIDR 5.7584924e-05
7,170 Lowering the Latency of Data Processing Pipelines Through FPGA based Hardware Acceleration 2020 VLDB 5.6840081e-05
Previous Page 1 / 1 Next

Semantically Similar Papers