DBScholar

Back to papers

MRShare: Sharing Across Multiple Queries in MapReduce

Summary: MRShare merges related MapReduce jobs into groups and runs them as a single query to share work across jobs. Using a MapReduce-specific cost model, it derives an optimal grouping to maximize overlap, with a Hadoop prototype demonstrating substantial cost savings. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
he2db0942238feff1
Venue
VLDB
Year
2010
Pagerank
0.00013642066
Overall Rank
823 | 94.48%
DOI
10.14778/1920841.1920905

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{nykiel_vldb10,
        title = {{MRShare: Sharing Across Multiple Queries in MapReduce}},
        author = {Nykiel, Tomasz and Potamias, Michalis and Mishra, Chaitanya and Kollios, George and Koudas, Nick},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {494--505},
        doi = {10.14778/1920841.1920905},
        url = {https://doi.org/10.14778/1920841.1920905},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 27 of 27 citing papers.

Rank Citing Paper Year Venue Pagerank
754 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014231311
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,747 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7303647e-05
1,925 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3643089e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3386546e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2782871e-05
2,682 Accelerating Recommendation System Training by Leveraging Popular Choices 2022 VLDB 8.1384948e-05
2,689 Major Technical Advancements in Apache Hive 2014 SIGMOD 8.1227087e-05
2,724 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.090674e-05
2,932 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8368679e-05
3,034 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7301387e-05
3,217 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5199872e-05
3,544 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2108612e-05
3,841 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.98877e-05
5,091 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 6.2775613e-05
5,278 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.1976381e-05
5,836 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9732317e-05
6,076 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8918218e-05
6,222 Materialization and Reuse Optimizations for Production Data Science Pipelines 2022 SIGMOD 5.8446676e-05
7,959 Distributed Outlier Detection using Compressive Sensing 2015 SIGMOD 5.4160954e-05
9,714 Hippo: Sharing Computations in Hyper-Parameter Optimization 2022 VLDB 5.1347788e-05
12,337 An Efficient MapReduce Cube Algorithm for Varied Data Distributions 2016 SIGMOD 4.9769913e-05
12,381 Parallel Evaluation of Multi-Semi-Joins 2016 VLDB 4.9769913e-05
12,453 Shared Execution of Recurring Workloads in MapReduce 2015 VLDB 4.9769913e-05
12,471 Anti-Combining for MapReduce 2014 SIGMOD 4.9769913e-05
12,565 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010515896
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050475202
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049821554
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
75 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.0003702496
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031099083
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028568843
155 Common Expression Analysis in Database Applications 1982 SIGMOD 0.0002851688
390 QPipe: A Simultaneously Pipelined Relational Query Engine 2005 SIGMOD 0.00019265472
651 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015123701
704 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014653079
762 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014134432
1,034 A Scalable, Predictable Join Operator for Highly Concurrent Data Warehouses 2009 VLDB 0.00012389548
1,132 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011893781
1,285 SQL/MapReduce: A practical approach to self-describing, polymorphic, and parallelizable user-defined functions 2009 VLDB 0.0001119616
3,314 Scheduling Shared Scans of Large Data Files 2008 VLDB 7.4349711e-05
Previous Page 1 / 1 Next

Semantically Similar Papers