DBScholar

Back to papers

MRShare: Sharing Across Multiple Queries in MapReduce

Summary: MRShare merges related MapReduce jobs into groups and runs them as a single query to share work across jobs. Using a MapReduce-specific cost model, it derives an optimal grouping to maximize overlap, with a Hadoop prototype demonstrating substantial cost savings. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
he2db0942238feff1
Venue
VLDB
Year
2010
Pagerank
0.00013648332
Overall Rank
823 | 94.47%
DOI
10.14778/1920841.1920905

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{nykiel_vldb10,
        title = {{MRShare: Sharing Across Multiple Queries in MapReduce}},
        author = {Nykiel, Tomasz and Potamias, Michalis and Mishra, Chaitanya and Kollios, George and Koudas, Nick},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {494--505},
        doi = {10.14778/1920841.1920905},
        url = {https://doi.org/10.14778/1920841.1920905},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 27 of 27 citing papers.

Rank Citing Paper Year Venue Pagerank
753 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014237583
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,745 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7343818e-05
1,924 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3687009e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3425785e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2821647e-05
2,682 Accelerating Recommendation System Training by Leveraging Popular Choices 2022 VLDB 8.1423493e-05
2,689 Major Technical Advancements in Apache Hive 2014 SIGMOD 8.1264718e-05
2,723 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.0944934e-05
2,931 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8405483e-05
3,033 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7337998e-05
3,216 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5234702e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
3,840 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.9920793e-05
5,088 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 6.2805079e-05
5,274 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.2005706e-05
5,834 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9759267e-05
6,074 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8946121e-05
6,217 Materialization and Reuse Optimizations for Production Data Science Pipelines 2022 SIGMOD 5.8474357e-05
7,955 Distributed Outlier Detection using Compressive Sensing 2015 SIGMOD 5.4185826e-05
9,709 Hippo: Sharing Computations in Hyper-Parameter Optimization 2022 VLDB 5.1372107e-05
12,331 An Efficient MapReduce Cube Algorithm for Varied Data Distributions 2016 SIGMOD 4.9793485e-05
12,375 Parallel Evaluation of Multi-Semi-Joins 2016 VLDB 4.9793485e-05
12,447 Shared Execution of Recurring Workloads in MapReduce 2015 VLDB 4.9793485e-05
12,465 Anti-Combining for MapReduce 2014 SIGMOD 4.9793485e-05
12,559 Mosquito: Another One Bites the Data Upload STream 2013 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
75 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.0003704106
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028579704
155 Common Expression Analysis in Database Applications 1982 SIGMOD 0.00028527932
389 QPipe: A Simultaneously Pipelined Relational Query Engine 2005 SIGMOD 0.00019269777
651 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015130782
705 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014657491
762 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014140446
1,034 A Scalable, Predictable Join Operator for Highly Concurrent Data Warehouses 2009 VLDB 0.00012394538
1,131 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.0001189909
1,285 SQL/MapReduce: A practical approach to self-describing, polymorphic, and parallelizable user-defined functions 2009 VLDB 0.00011201377
3,313 Scheduling Shared Scans of Large Data Files 2008 VLDB 7.4383267e-05
Previous Page 1 / 1 Next

Semantically Similar Papers