DBScholar

Back to papers

Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience

Summary: Pig provides SQL-like data manipulation on MapReduce by building explicit dataflows interleaved with UDFs, compiled to Hadoop jobs. It discusses challenges and compares Pig's performance to hand-tuned MapReduce, showing productivity gains with modest overhead. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc0adf90ae67c937b
Venue
VLDB
Year
2009
Pagerank
0.00015130782
Overall Rank
651 | 95.63%
DOI
10.14778/1687553.1687577

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gates_vldb09,
        title = {{Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience}},
        author = {Gates, Alan F. and Natkovich, Olga and Chopra, Shubham and Kamath, Pradeep and Narayanamurthy, Shravan M. and Olston, Christopher and Reed, Benjamin and Srinivasan, Santhosh and Srivastava, Utkarsh},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1414--1425},
        doi = {10.14778/1687553.1687577},
        url = {https://doi.org/10.14778/1687553.1687577},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 33 of 33 citing papers.

Rank Citing Paper Year Venue Pagerank
325 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.00020964941
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
977 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012731794
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,080 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012131978
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010934347
1,837 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5350058e-05
1,924 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3687009e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3425785e-05
2,635 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.1965216e-05
3,216 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5234702e-05
3,636 Advanced Join Strategies for Large-Scale Distributed Computation 2014 VLDB 7.1471007e-05
3,788 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0195464e-05
3,823 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.0029055e-05
4,069 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.821366e-05
4,249 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7037533e-05
4,254 Nova: Continuous Pig/Hadoop Workflows 2011 SIGMOD 6.7011496e-05
4,771 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 VLDB 6.4230759e-05
5,058 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 6.2926774e-05
5,574 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 SIGMOD 6.0794039e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
5,828 Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture 2013 SIGMOD 5.9786171e-05
6,074 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8946121e-05
6,422 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.7901142e-05
7,308 Optimization for iterative queries on MapReduce 2014 VLDB 5.5572894e-05
8,004 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4089097e-05
9,307 SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory 2014 SIGMOD 5.197526e-05
9,693 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1399537e-05
9,816 Supporting Scalable Analytics with Latency Constraints 2015 VLDB 5.1257999e-05
12,465 Anti-Combining for MapReduce 2014 SIGMOD 4.9793485e-05
12,689 Resiliency-Aware Data Management 2011 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers