DBScholar

Back to papers

Pig Latin: A Not-So-Foreign Language for Data Processing

Summary: Pig Latin sits between SQL and MapReduce, enabling procedural analysts to express data flows without MapReduce coding. Pig compiles Pig Latin to Hadoop MapReduce plans and offers an integrated debugger; open-source under Apache Incubator with Yahoo-scale deployments. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4119
Venue
SIGMOD
Year
2008
Pagerank
0.0010686205
Overall Rank
6 | 99.97%
DOI
10.1145/1376616.1376726

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{olston_sigmod08,
        title = {{Pig Latin: A Not-So-Foreign Language for Data Processing}},
        author = {Olston, Christopher and Reed, Benjamin and Srivastava, Utkarsh and Kumar, Ravi and Tomkins, Andrew},
        series = {{SIGMOD} '08},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1376616.1376726},
        url = {https://dl.acm.org/doi/10.1145/1376616.1376726},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 153 citing papers.

Rank Citing Paper Year Venue Pagerank
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,622 Good to the Last Bit: Data-Driven Encoding with CodecDB 2021 SIGMOD 6.1461066e-05
5,748 Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture 2013 SIGMOD 6.1014813e-05
5,785 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 6.0892672e-05
6,147 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.9601328e-05
6,181 Just-In-Time Data Virtualization: Lightweight Data Management with ViDa 2015 CIDR 5.9494838e-05
6,377 Scalable Querying of Nested Data 2021 VLDB 5.8931544e-05
6,382 The Era of Big Spatial Data 2017 VLDB 5.8915267e-05
6,430 REEF: Retainable Evaluator Execution Framework 2015 SIGMOD 5.880521e-05
6,557 Semistructured Models, Queries and Algebras in the Big Data Era 2016 SIGMOD 5.8401127e-05
6,714 Towards Unified Ad-hoc Data Processing 2014 SIGMOD 5.7959651e-05
6,736 Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads 2013 VLDB 5.787547e-05
6,822 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 5.7627651e-05
6,889 An Algebraic Approach for Data-Centric Scientific Workflows 2011 VLDB 5.7455532e-05
6,896 Wide Table Layout Optimization based on Column Ordering and Duplication 2017 SIGMOD 5.7440611e-05
7,062 BSMA: A Benchmark for Analytical Queries over Social Media Data 2014 VLDB 5.7133535e-05
7,168 Optimization for iterative queries on MapReduce 2014 VLDB 5.6841364e-05
7,238 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.664642e-05
7,243 Online Expansion of Large-scale Data Warehouses 2011 VLDB 5.6640533e-05
7,551 Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams 2022 VLDB 5.6006414e-05
7,689 Oracle In-Database Hadoop: When MapReduce Meets RDBMS 2012 SIGMOD 5.5669484e-05
7,798 A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data 2017 VLDB 5.542286e-05
7,965 Emerging Trends in the Enterprise Data Analytics: Connecting Hadoop and DB2 Warehouse 2011 SIGMOD 5.5174991e-05
8,272 Shasta: Interactive Reporting At Scale 2016 SIGMOD 5.4574671e-05
8,278 Building Community-Centric Information Exploration Applications on Social Content Sites 2009 SIGMOD 5.4574671e-05
8,429 Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis 2015 CIDR 5.4263087e-05
8,455 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.4226e-05
8,668 Toward Progress Indicators on Steroids for Big Data Systems 2013 CIDR 5.3882971e-05
8,712 Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler 2017 SIGMOD 5.3779634e-05
8,955 From SPARQL to MapReduce: The Journey Using a Nested TripleGroup Algebra 2011 VLDB 5.3457324e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
9,079 QMapper for Smart Grid: Migrating SQL-based Application to Hive 2015 SIGMOD 5.3251649e-05
9,137 SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory 2014 SIGMOD 5.3162511e-05
9,274 DataGarage: Warehousing Massive Performance Data on Commodity Servers 2010 VLDB 5.2941406e-05
9,497 Rank Join Queries in NoSQL Databases 2014 VLDB 5.2608378e-05
9,509 Versatile Optimization of UDF-heavy Data Flows with Sofa 2014 SIGMOD 5.2581466e-05
9,661 PAXQuery: Parallel Analytical XML Processing 2015 SIGMOD 5.2409648e-05
9,748 Graft: A Debugging Tool For Apache Giraph 2015 SIGMOD 5.227679e-05
11,399 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark 2023 SIGMOD 5.093636e-05
11,414 Udon: Efficient Debugging of User-Defined Functions in Big Data Systems with Line-by-Line Control 2023 SIGMOD 5.093636e-05
11,885 Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology 2019 VLDB 5.093636e-05
12,032 Logical Aspects of Massively Parallel and Distributed Systems 2016 PODS 5.093636e-05
12,060 dmapply: A functional primitive to express distributed machine learning algorithms in R 2016 VLDB 5.093636e-05
12,082 Parallel Evaluation of Multi-Semi-Joins 2016 VLDB 5.093636e-05
12,090 Let's Rethink Join Optimization in Distributed Systems 2015 CIDR 5.093636e-05
12,115 A Demonstration of Rubato DB: A Highly Scalable NewSQL Database System for OLTP and Big Data Applications 2015 SIGMOD 5.093636e-05
12,118 ShareInsights - An Unified Approach to Full-stack Data Processing 2015 SIGMOD 5.093636e-05
12,174 Anti-Combining for MapReduce 2014 SIGMOD 5.093636e-05
12,306 Declarative Error Management for Robust Data-Intensive Applications 2012 SIGMOD 5.093636e-05
12,322 ReStore: Reusing Results of MapReduce Jobs in Pig 2012 SIGMOD 5.093636e-05
Previous Page 3 / 4 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
54 On Random Sampling over Joins 1999 SIGMOD 0.00040810225
72 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.00037695852
Previous Page 1 / 1 Next

Semantically Similar Papers