M3R: Increased Performance for In-Memory Hadoop Jobs
Summary: M3R is an in-memory Hadoop MapReduce engine for online analytics on memory-resident clusters, sacrificing resilience for speed. Unchanged HMR execution (Pig/Jaql/SystemML, BigSheets) with large speedups (≈45× for sparse mat-vec) and API extensions that accelerate workloads without altering Hadoop semantics. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Avraham Shinnar (IBM)
- 2. David Cunningham (IBM)
- 3. Benjamin Herta (IBM)
- 4. Vijay Saraswat (IBM)
BibTeX Citation
@article{shinnar_vldb12,
title = {{M3R: Increased Performance for In-Memory Hadoop Jobs}},
author = {Shinnar, Avraham and Cunningham, David and Herta, Benjamin and Saraswat, Vijay},
journal = {PVLDB},
series = {{VLDB} '12},
volume = {5},
number = {12},
pages = {1736--1747},
doi = {10.14778/2367502.2367517},
url = {https://doi.org/10.14778/2367502.2367517},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,573 | Minimal MapReduce Algorithms | 2013 | SIGMOD | 8.2821647e-05 |
| 2,739 | General Incremental Sliding-Window Aggregation | 2015 | VLDB | 8.0721161e-05 |
| 5,408 | Lifetime-Based Memory Management for Distributed Data Processing Systems | 2016 | VLDB | 6.1415736e-05 |
| 8,100 | Hone: “Scaling Down” Hadoop on Shared-Memory Systems | 2013 | VLDB | 5.3942942e-05 |
| 12,461 | Palette: Enabling Scalable Analytics for Big-Memory, Multicore Machines | 2014 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.001052036 |
| 31 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00049839909 |
| 384 | HaLoop: Efficient Iterative Data Processing on Large Clusters | 2010 | VLDB | 0.0001948031 |
| 1,037 | Jaql: A Scripting Language for Large Scale Semistructured Data Analysis | 2011 | VLDB | 0.00012376819 |
| 1,302 | Apache Hadoop Goes Realtime at Facebook | 2011 | SIGMOD | 0.00011111126 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,100 | Hone: “Scaling Down” Hadoop on Shared-Memory Systems | 2013 | VLDB |
| 2 | 3,216 | Multi-Query Optimization in MapReduce Framework | 2014 | VLDB |
| 3 | 13,857 | M3: Scaling Up Machine Learning via Memory Mapping | 2016 | SIGMOD |
| 4 | 673 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 5 | 1,924 | ReStore: Reusing Results of MapReduce Jobs | 2012 | VLDB |
| 6 | 1,466 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 7 | 2,282 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 8 | 75 | Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters | 2007 | SIGMOD |
| 9 | 2,361 | Online Aggregation and Continuous Query support in MapReduce | 2010 | SIGMOD |
| 10 | 7,632 | MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs | 2014 | VLDB |