Bridging the Gap Between HPC and Big Data Frameworks
Summary: MPI-Spark integration to offload compute from Spark into MPI, preserving Spark's fault tolerance and ecosystem. Analyzes four distributed graph/ML workloads; shows 3.1-17.7x speedups over native Spark with overheads included, enabling reuse of MPI libraries in Spark with minimal effort. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Anderson (Intel)
- 2. Shaden Smith (University of Minnesota)
- 3. Narayanan Sundaram (Intel)
- 4. Mihai Capotă (Intel)
- 5. Zheguang Zhao (Brown University)
- 6. Subramanya Dulloor (Intel)
- 7. Nadathur Satish (Intel)
- 8. Theodore L. Willke (Intel)
BibTeX Citation
@article{anderson_vldb17,
title = {{Bridging the Gap Between HPC and Big Data Frameworks}},
author = {Anderson, Michael and Smith, Shaden and Sundaram, Narayanan and Capotă, Mihai and Zhao, Zheguang and Dulloor, Subramanya and Satish, Nadathur and Willke, Theodore L.},
journal = {PVLDB},
series = {{VLDB} '17},
volume = {10},
number = {8},
pages = {901--912},
doi = {10.14778/3093688.3093693},
url = {https://doi.org/10.14778/3093688.3093693},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,959 | Datalog with First-Class Facts | 2025 | VLDB | 5.1879626e-05 |
| 10,076 | Chukonu: A Fully-Featured High-Performance Big Data Framework that Integrates a Native Compute Engine into Spark | 2022 | VLDB | 5.1613298e-05 |
| 11,756 | Approximate Pattern Matching in Massive Graphs with Precision and Recall Guarantees | 2020 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,654 | Navigating the Maze of Graph Analytics Frameworks using Massive Graph Datasets | 2014 | SIGMOD | 0.00010104703 |
| 1,966 | GraphMat: High performance graph analytics made productive | 2015 | VLDB | 9.3844743e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,615 | A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning | 2024 | VLDB |
| 2 | 13,550 | Trends and Challenges in Big Data Processing | 2016 | VLDB |
| 3 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 4 | 8,738 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB |
| 5 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 6 | 7,344 | Building the Enterprise Fabric for Big Data with Vertica and Spark Integration | 2016 | SIGMOD |
| 7 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |
| 8 | 8,104 | Enabling Transparent Acceleration of Big Data Frameworks Using Heterogeneous Hardware | 2022 | VLDB |
| 9 | 4,208 | Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics | 2015 | VLDB |
| 10 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |