DBScholar

Back to papers

Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience

Summary: Pig provides SQL-like data manipulation on MapReduce by building explicit dataflows interleaved with UDFs, compiled to Hadoop jobs. It discusses challenges and compares Pig's performance to hand-tuned MapReduce, showing productivity gains with modest overhead. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
10024
Venue
VLDB
Year
2009
Pagerank
0.00015395331
Overall Rank
642 | 95.60%
DOI
10.14778/1687553.1687577

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gates_vldb09,
        title = {{Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience}},
        author = {Gates, Alan F. and Natkovich, Olga and Chopra, Shubham and Kamath, Pradeep and Narayanamurthy, Shravan M. and Olston, Christopher and Reed, Benjamin and Srinivasan, Santhosh and Srivastava, Utkarsh},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1414--1425},
        doi = {10.14778/1687553.1687577},
        url = {https://doi.org/10.14778/1687553.1687577},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 33 of 33 citing papers.

Rank Citing Paper Year Venue Pagerank
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
361 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.00020138717
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
954 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012997301
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,067 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012327784
1,319 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00011175005
1,791 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.7470504e-05
1,883 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.5421713e-05
2,486 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.5143189e-05
2,594 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.3646367e-05
3,163 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.6784171e-05
3,578 Advanced Join Strategies for Large-Scale Distributed Computation 2014 VLDB 7.2899943e-05
3,714 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.1764857e-05
3,749 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.1539955e-05
4,169 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.8561463e-05
4,175 Nova: Continuous Pig/Hadoop Workflows 2011 SIGMOD 6.8511274e-05
4,481 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.6754521e-05
4,693 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 VLDB 6.560102e-05
5,388 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 6.2362811e-05
5,442 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 SIGMOD 6.2184416e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,748 Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture 2013 SIGMOD 6.1014813e-05
6,044 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.995949e-05
6,312 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.9170938e-05
7,168 Optimization for iterative queries on MapReduce 2014 VLDB 5.6841364e-05
8,615 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4005602e-05
9,137 SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory 2014 SIGMOD 5.3162511e-05
9,510 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.2576928e-05
9,640 Supporting Scalable Analytics with Latency Constraints 2015 VLDB 5.2434488e-05
12,174 Anti-Combining for MapReduce 2014 SIGMOD 5.093636e-05
12,398 Resiliency-Aware Data Management 2011 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers