Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph Processing
Summary: Clementi is a multi-FPGA graph-processing framework with fine-grained compute and cross-FPGA pipelines to tackle irregular data transfer and workload skew. A precise execution-time model enables a scheduler that balances load and overlaps communication with computation, delivering up to 8.75x speedups and near-linear scalability. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Feng Yu (National University of Singapore)
- 2. Hongshi Tan (National University of Singapore)
- 3. Xinyu Chen (Hong Kong University of Science and Technology)
- 4. Yao Chen (National University of Singapore)
- 5. Bingsheng He (National University of Singapore)
- 6. Weng-fai Wong (National University of Singapore)
BibTeX Citation
@inproceedings{yu_sigmod25,
title = {{Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph Processing}},
author = {Yu, Feng and Tan, Hongshi and Chen, Xinyu and Chen, Yao and He, Bingsheng and Wong, Weng-fai},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725275},
url = {https://dl.acm.org/doi/10.1145/3725275},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3 | Pregel: A System for Large-Scale Graph Processing | 2010 | SIGMOD | 0.0012250108 |
| 20 | Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud | 2012 | VLDB | 0.00056944564 |
| 1,511 | Speedup Graph Processing by Graph Ordering | 2016 | SIGMOD | 0.00010538011 |
| 1,654 | Navigating the Maze of Graph Analytics Frameworks using Massive Graph Datasets | 2014 | SIGMOD | 0.00010104703 |
| 3,303 | A Distributed Multi-GPU System for Fast Graph Processing | 2018 | VLDB | 7.5405596e-05 |
| 5,942 | Cardinality Estimation of Subgraph Matching: A Filtering-Sampling Approach | 2024 | VLDB | 6.0334209e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,085 | Start Late or Finish Early: A Distributed Graph Processing System with Redundancy Reduction | 2019 | VLDB |
| 2 | 10,950 | Efficient Graph Data Access for Out-of-Memory GPU Streaming Graph Processing | 2025 | VLDB |
| 3 | 10,236 | FaaSBoard: Efficient Graph Processing with a Disaggregated Architecture on Serverless Services | 2026 | SIGMOD |
| 4 | 9,605 | Nezha: An Efficient Distributed Graph Processing System on Heterogeneous Hardware | 2025 | SIGMOD |
| 5 | 7,380 | MiniGraph: Querying Big Graphs with a Single Machine | 2023 | VLDB |
| 6 | 7,170 | Lowering the Latency of Data Processing Pipelines Through FPGA based Hardware Acceleration | 2020 | VLDB |
| 7 | 922 | Data Processing on FPGAs | 2009 | VLDB |
| 8 | 5,480 | EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal in GPUs | 2021 | VLDB |
| 9 | 4,157 | GPU-based Graph Traversal on Compressed Graphs | 2019 | SIGMOD |
| 10 | 5,075 | CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processor | 2024 | VLDB |