Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications
Summary: Casper translates Java programs into MapReduce-style implementations for Hadoop, Spark, and Flink. It uses program synthesis to infer a MapReduce summary, verified by a theorem prover, then emits executable code; benchmarks show up to 48.2x speedups. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Maaz Bin Safeer Ahmad (University of Washington)
- 2. Alvin Cheung (University of Washington)
BibTeX Citation
@inproceedings{ahmad_sigmod18,
title = {{Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications}},
author = {Ahmad, Maaz Bin Safeer and Cheung, Alvin},
series = {{SIGMOD} '18},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3183713.3196891},
url = {https://dl.acm.org/doi/10.1145/3183713.3196891},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,293 | Extending Relational Query Processing with ML Inference | 2020 | CIDR | 8.7949378e-05 |
| 3,992 | Aggify: Lifting the Curse of Cursor Loops using Custom Aggregates | 2020 | SIGMOD | 6.9695338e-05 |
| 5,265 | New Directions in Cloud Programming | 2021 | CIDR | 6.2941668e-05 |
| 8,465 | Predicate Pushdown for Data Science Pipelines | 2023 | SIGMOD | 5.4194578e-05 |
| 8,738 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB | 5.3766157e-05 |
| 11,499 | Towards Auto-Generated Data Systems | 2023 | VLDB | 5.093636e-05 |
| 11,711 | TraNCE: Transforming Nested Collections Efficiently | 2021 | VLDB | 5.093636e-05 |
| 11,739 | View-Driven Optimization of Database-Backed Web Applications | 2020 | CIDR | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1 | Access Path Selection in a Relational Database Management System | 1979 | SIGMOD | 0.0024089429 |
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB | 0.0001888524 |
| 534 | Building Efficient Query Engines in a High-Level Language | 2014 | VLDB | 0.00017046514 |
| 1,229 | Weld: A Common Runtime for High Performance Data Analytics | 2017 | CIDR | 0.00011578425 |
| 2,717 | Implicit Parallelism through Deep Language Embedding | 2015 | SIGMOD | 8.2102313e-05 |
| 2,939 | Extracting Equivalent SQL from Imperative Code in Database Applications | 2016 | SIGMOD | 7.9395908e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 44 | A Comparison of Approaches to Large-Scale Data Analysis | 2009 | SIGMOD |
| 2 | 9,079 | QMapper for Smart Grid: Migrating SQL-based Application to Hive | 2015 | SIGMOD |
| 3 | 12,131 | FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data | 2015 | VLDB |
| 4 | 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 5 | 1,054 | Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads | 2012 | VLDB |
| 6 | 2,717 | Implicit Parallelism through Deep Language Embedding | 2015 | SIGMOD |
| 7 | 1,257 | Automatic Optimization for MapReduce Programs | 2011 | VLDB |
| 8 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 9 | 735 | Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs | 2011 | VLDB |
| 10 | 13,531 | Optimizing Data-Intensive Applications Automatically By Leveraging Parallel Data Processing Frameworks | 2017 | SIGMOD |