DBScholar

Back to papers

MRShare: Sharing Across Multiple Queries in MapReduce

Summary: MRShare merges related MapReduce jobs into groups and runs them as a single query to share work across jobs. Using a MapReduce-specific cost model, it derives an optimal grouping to maximize overlap, with a Hadoop prototype demonstrating substantial cost savings. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
10290
Venue
VLDB
Year
2010
Pagerank
0.00013899943
Overall Rank
803 | 94.50%
DOI
10.14778/1920841.1920905

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{nykiel_vldb10,
        title = {{MRShare: Sharing Across Multiple Queries in MapReduce}},
        author = {Nykiel, Tomasz and Potamias, Michalis and Mishra, Chaitanya and Kollios, George and Koudas, Nick},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {494--505},
        doi = {10.14778/1920841.1920905},
        url = {https://doi.org/10.14778/1920841.1920905},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 27 of 27 citing papers.

Rank Citing Paper Year Venue Pagerank
735 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014522606
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,765 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.8079546e-05
1,883 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.5421713e-05
2,486 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.5143189e-05
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
2,681 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.2632778e-05
2,688 Accelerating Recommendation System Training by Leveraging Popular Choices 2022 VLDB 8.2564305e-05
2,706 Major Technical Advancements in Apache Hive 2014 SIGMOD 8.2287564e-05
2,887 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.9952432e-05
3,117 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7382559e-05
3,163 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.6784171e-05
3,605 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2640711e-05
3,783 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 7.1298683e-05
5,156 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.3410921e-05
5,720 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 6.1092082e-05
6,044 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.995949e-05
6,309 Materialization and Reuse Optimizations for Production Data Science Pipelines 2022 SIGMOD 5.9189554e-05
7,673 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 5.5705944e-05
7,795 Distributed Outlier Detection using Compressive Sensing 2015 SIGMOD 5.5429229e-05
9,529 Hippo: Sharing Computations in Hyper-Parameter Optimization 2022 VLDB 5.2545475e-05
12,036 An Efficient MapReduce Cube Algorithm for Varied Data Distributions 2016 SIGMOD 5.093636e-05
12,082 Parallel Evaluation of Multi-Semi-Joins 2016 VLDB 5.093636e-05
12,156 Shared Execution of Recurring Workloads in MapReduce 2015 VLDB 5.093636e-05
12,174 Anti-Combining for MapReduce 2014 SIGMOD 5.093636e-05
12,268 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00051174276
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
72 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.00037695852
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031680027
153 Common Expression Analysis in Database Applications 1982 SIGMOD 0.00029032276
155 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028713176
383 QPipe: A Simultaneously Pipelined Relational Query Engine 2005 SIGMOD 0.00019520728
642 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015395331
696 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014891322
761 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014254351
1,022 A Scalable, Predictable Join Operator for Highly Concurrent Data Warehouses 2009 VLDB 0.00012602841
1,154 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011934202
1,264 SQL/MapReduce: A practical approach to self-describing, polymorphic, and parallelizable user-defined functions 2009 VLDB 0.00011416393
3,270 Scheduling Shared Scans of Large Data Files 2008 VLDB 7.5796338e-05
Previous Page 1 / 1 Next

Semantically Similar Papers