sPCA: Scalable Principal Component Analysis for Big Data on Distributed Platforms
Summary: Introduces sPCA, a scalable PCA optimized for distributed big-data platforms. Leverages sparse matrix ops, minimizes intermediates, and is implemented on MapReduce and Spark; outperforms Mahout-PCA and MLlib-PCA in accuracy, speed, and data-shuffle. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tarek Elgamal (Qatar Computing Research Institute)
- 2. Maysam Yabandeh (Twitter)
- 3. Ashraf Aboulnaga (Qatar Computing Research Institute)
- 4. Waleed Mustafa (NTG Clarity)
- 5. Mohamed Hefeeda (Qatar Computing Research Institute)
BibTeX Citation
@inproceedings{elgamal_sigmod15,
title = {{sPCA: Scalable Principal Component Analysis for Big Data on Distributed Platforms}},
author = {Elgamal, Tarek and Yabandeh, Maysam and Aboulnaga, Ashraf and Mustafa, Waleed and Hefeeda, Mohamed},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2751520},
url = {https://dl.acm.org/doi/10.1145/2723372.2751520},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,251 | QCore: Data-Efficient, On-Device Continual Calibration for Quantized Models | 2024 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 372 | HaLoop: Efficient Iterative Data Processing on Large Clusters | 2010 | VLDB | 0.0001981521 |
| 532 | MLbase: A Distributed Machine-learning System | 2013 | CIDR | 0.00017072641 |
| 1,742 | ArrayStore: A Storage Manager for Complex Parallel Array Processing | 2011 | SIGMOD | 9.8748669e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 13,544 | M3: Scaling Up Machine Learning via Memory Mapping | 2016 | SIGMOD |
| 2 | 2,539 | Minimal MapReduce Algorithms | 2013 | SIGMOD |
| 3 | 6,850 | Bridging the Gap Between HPC and Big Data Frameworks | 2017 | VLDB |
| 4 | 4,208 | Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics | 2015 | VLDB |
| 5 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 6 | 11,621 | Scalable Robust Graph Embedding with Spark | 2022 | VLDB |
| 7 | 3,008 | Scalable Big Graph Processing in MapReduce | 2014 | SIGMOD |
| 8 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |
| 9 | 12,036 | An Efficient MapReduce Cube Algorithm for Varied Data Distributions | 2016 | SIGMOD |
| 10 | 8,268 | Efficient Matrix Sketching over Distributed Data | 2017 | PODS |