Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications
Summary: Casper translates Java programs into MapReduce-style implementations for Hadoop, Spark, and Flink. It uses program synthesis to infer a MapReduce summary, verified by a theorem prover, then emits executable code; benchmarks show up to 48.2x speedups. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Maaz Bin Safeer Ahmad (University of Washington)
- 2. Alvin Cheung (University of Washington)
BibTeX Citation
@inproceedings{ahmad_sigmod18,
title = {{Automatically Leveraging MapReduce Frameworks for Data-Intensive Applications}},
author = {Ahmad, Maaz Bin Safeer and Cheung, Alvin},
series = {{SIGMOD} '18},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3183713.3196891},
url = {https://dl.acm.org/doi/10.1145/3183713.3196891},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,182 | Extending Relational Query Processing with ML Inference | 2020 | CIDR | 8.8982998e-05 |
| 4,061 | Aggify: Lifting the Curse of Cursor Loops using Custom Aggregates | 2020 | SIGMOD | 6.824809e-05 |
| 5,380 | New Directions in Cloud Programming | 2021 | CIDR | 6.1551126e-05 |
| 8,153 | Predicate Pushdown for Data Science Pipelines | 2023 | SIGMOD | 5.3891786e-05 |
| 8,900 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB | 5.2559789e-05 |
| 11,808 | Towards Auto-Generated Data Systems | 2023 | VLDB | 4.9793485e-05 |
| 12,015 | TraNCE: Transforming Nested Collections Efficiently | 2021 | VLDB | 4.9793485e-05 |
| 12,042 | View-Driven Optimization of Database-Backed Web Applications | 2020 | CIDR | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1 | Access Path Selection in a Relational Database Management System | 1979 | SIGMOD | 0.0023947656 |
| 23 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00055406774 |
| 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB | 0.0001865959 |
| 495 | Building Efficient Query Engines in a High-Level Language | 2014 | VLDB | 0.00017370758 |
| 1,193 | Weld: A Common Runtime for High Performance Data Analytics | 2017 | CIDR | 0.0001158809 |
| 2,754 | Implicit Parallelism through Deep Language Embedding | 2015 | SIGMOD | 8.0534972e-05 |
| 2,993 | Extracting Equivalent SQL from Imperative Code in Database Applications | 2016 | SIGMOD | 7.7721951e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 43 | A Comparison of Approaches to Large-Scale Data Analysis | 2009 | SIGMOD |
| 2 | 9,254 | QMapper for Smart Grid: Migrating SQL-based Application to Hive | 2015 | SIGMOD |
| 3 | 12,424 | FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data | 2015 | VLDB |
| 4 | 673 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 5 | 1,072 | Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads | 2012 | VLDB |
| 6 | 2,754 | Implicit Parallelism through Deep Language Embedding | 2015 | SIGMOD |
| 7 | 1,281 | Automatic Optimization for MapReduce Programs | 2011 | VLDB |
| 8 | 2,282 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 9 | 753 | Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs | 2011 | VLDB |
| 10 | 13,844 | Optimizing Data-Intensive Applications Automatically By Leveraging Parallel Data Processing Frameworks | 2017 | SIGMOD |