Cumulon: Optimizing Statistical Data Analysis in the Cloud
Summary: Cumulon: cloud-native system for rapid matrix analytics development and deployment. Automatic optimization across operators, parameters, and hardware provisioning under time/budget constraints; implemented atop Hadoop/HDFS to avoid MapReduce limits. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Botong Huang (Duke University)
- 2. Shivnath Babu (Duke University)
- 3. Jun Yang (Duke University)
BibTeX Citation
@inproceedings{huang_sigmod13,
title = {{Cumulon: Optimizing Statistical Data Analysis in the Cloud}},
author = {Huang, Botong and Babu, Shivnath and Yang, Jun},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2465273},
url = {https://dl.acm.org/doi/10.1145/2463676.2465273},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010686205 |
| 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00050111008 |
| 106 | The MADlib Analytics Library or MAD Skills, the SQL | 2012 | VLDB | 0.00033539462 |
| 155 | MAD Skills: New Analysis Practices for Big Data | 2009 | VLDB | 0.00028713176 |
| 239 | Overview of SciDB: Large Scale Array Storage, Processing and Analysis | 2010 | SIGMOD | 0.00023674329 |
| 372 | HaLoop: Efficient Iterative Data Processing on Large Clusters | 2010 | VLDB | 0.0001981521 |
| 735 | Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs | 2011 | VLDB | 0.00014522606 |
| 1,039 | RIOT: I/O-Efficient Numerical Computing without SQL | 2009 | CIDR | 0.00012474977 |
| 1,333 | Ricardo: Integrating R and Hadoop | 2010 | SIGMOD | 0.0001112858 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 923 | Starfish: A Self-tuning System for Big Data Analytics | 2011 | CIDR |
| 2 | 3,671 | Exploiting MapReduce-based Similarity Joins | 2012 | SIGMOD |
| 3 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 4 | 2,642 | CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop | 2011 | VLDB |
| 5 | 2,921 | Clustera: An Integrated Computation And Data Management System | 2008 | VLDB |
| 6 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 7 | 13,626 | Data Mining Algorithms as a Service in the Cloud: Exploiting Relational Database Systems | 2013 | SIGMOD |
| 8 | 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 9 | 13,540 | Cümülön-D: Data Analytics in a Dynamic Spot Market | 2017 | VLDB |
| 10 | 10,099 | Cumulon: Matrix-Based Data Analytics in the Cloud with Spot Instances | 2016 | VLDB |