DBScholar

Back to papers

Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience

Summary: Pig provides SQL-like data manipulation on MapReduce by building explicit dataflows interleaved with UDFs, compiled to Hadoop jobs. It discusses challenges and compares Pig's performance to hand-tuned MapReduce, showing productivity gains with modest overhead. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc0adf90ae67c937b
Venue
VLDB
Year
2009
Pagerank
0.00015123701
Overall Rank
651 | 95.63%
DOI
10.14778/1687553.1687577

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gates_vldb09,
        title = {{Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience}},
        author = {Gates, Alan F. and Natkovich, Olga and Chopra, Shubham and Kamath, Pradeep and Narayanamurthy, Shravan M. and Olston, Christopher and Reed, Benjamin and Srinivasan, Santhosh and Srivastava, Utkarsh},
        journal = {PVLDB},
        series = {{VLDB} '09},
        volume = {2},
        number = {2},
        pages = {1414--1425},
        doi = {10.14778/1687553.1687577},
        url = {https://doi.org/10.14778/1687553.1687577},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 33 of 33 citing papers.

Rank Citing Paper Year Venue Pagerank
325 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing 2015 VLDB 0.0002095522
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020001237
675 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00014880686
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013642066
978 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012725823
1,073 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012163258
1,081 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012126286
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010929229
1,838 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5305061e-05
1,925 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3643089e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3386546e-05
2,636 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.1926426e-05
3,217 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5199872e-05
3,637 Advanced Join Strategies for Large-Scale Distributed Computation 2014 VLDB 7.1437959e-05
3,790 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0162259e-05
3,824 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 6.9996014e-05
4,069 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.8184364e-05
4,250 The Unified Logging Infrastructure for Data Analytics at Twitter 2012 VLDB 6.7005978e-05
4,255 Nova: Continuous Pig/Hadoop Workflows 2011 SIGMOD 6.6979824e-05
4,774 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 VLDB 6.4201418e-05
5,062 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 6.2896995e-05
5,575 Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows 2011 SIGMOD 6.0765269e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
5,830 Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture 2013 SIGMOD 5.9757869e-05
6,076 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8918218e-05
6,424 A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses 2011 SIGMOD 5.7873769e-05
7,311 Optimization for iterative queries on MapReduce 2014 VLDB 5.5546605e-05
8,009 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4063491e-05
9,316 SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory 2014 SIGMOD 5.1951189e-05
9,699 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1375205e-05
9,823 Supporting Scalable Analytics with Latency Constraints 2015 VLDB 5.1233734e-05
12,471 Anti-Combining for MapReduce 2014 SIGMOD 4.9769913e-05
12,695 Resiliency-Aware Data Management 2011 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers