M3R: Increased Performance for In-Memory Hadoop Jobs
Summary: M3R is an in-memory Hadoop MapReduce engine for online analytics on memory-resident clusters, sacrificing resilience for speed. Unchanged HMR execution (Pig/Jaql/SystemML, BigSheets) with large speedups (≈45× for sparse mat-vec) and API extensions that accelerate workloads without altering Hadoop semantics. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Avraham Shinnar (IBM)
- 2. David Cunningham (IBM)
- 3. Benjamin Herta (IBM)
- 4. Vijay Saraswat (IBM)
BibTeX Citation
@article{shinnar_vldb12,
title = {{M3R: Increased Performance for In-Memory Hadoop Jobs}},
author = {Shinnar, Avraham and Cunningham, David and Herta, Benjamin and Saraswat, Vijay},
journal = {PVLDB},
series = {{VLDB} '12},
volume = {5},
number = {12},
pages = {1736--1747},
doi = {10.14778/2367502.2367517},
url = {https://doi.org/10.14778/2367502.2367517},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,539 | Minimal MapReduce Algorithms | 2013 | SIGMOD | 8.4526595e-05 |
| 2,802 | General Incremental Sliding-Window Aggregation | 2015 | VLDB | 8.1093063e-05 |
| 5,476 | Lifetime-Based Memory Management for Distributed Data Processing Systems | 2016 | VLDB | 6.2052672e-05 |
| 7,932 | Hone: “Scaling Down” Hadoop on Shared-Memory Systems | 2013 | VLDB | 5.5181056e-05 |
| 12,170 | Palette: Enabling Scalable Analytics for Big-Memory, Multicore Machines | 2014 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010686205 |
| 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00050111008 |
| 372 | HaLoop: Efficient Iterative Data Processing on Large Clusters | 2010 | VLDB | 0.0001981521 |
| 1,021 | Jaql: A Scripting Language for Large Scale Semistructured Data Analysis | 2011 | VLDB | 0.00012606673 |
| 1,324 | Apache Hadoop Goes Realtime at Facebook | 2011 | SIGMOD | 0.00011149314 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,932 | Hone: “Scaling Down” Hadoop on Shared-Memory Systems | 2013 | VLDB |
| 2 | 3,163 | Multi-Query Optimization in MapReduce Framework | 2014 | VLDB |
| 3 | 13,544 | M3: Scaling Up Machine Learning via Memory Mapping | 2016 | SIGMOD |
| 4 | 1,883 | ReStore: Reusing Results of MapReduce Jobs | 2012 | VLDB |
| 5 | 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 6 | 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 7 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 8 | 72 | Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters | 2007 | SIGMOD |
| 9 | 2,312 | Online Aggregation and Continuous Query support in MapReduce | 2010 | SIGMOD |
| 10 | 7,546 | MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs | 2014 | VLDB |