DBScholar

Back to papers

SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets

Summary: SCOPE is a declarative, extensible scripting language for data analysis on clusters, hiding parallelism from users. It provides SQL-like modeling with joins and aggregates, plus user-defined operators (extractors, processors, reducers, combiners), nesting, and stepwise plans compiled into parallel execution. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9943
Venue
VLDB
Year
2008
Pagerank
0.00051174276
Overall Rank
30 | 99.80%
DOI
10.14778/1454159.1454224

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chaiken_vldb08,
        title = {{SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets}},
        author = {Chaiken, Ronnie and Jenkins, Bob and Larson, Per-Åke and Ramsey, Bill and Shakib, Darren and Weaver, Simon and Zhou, Jingren},
        journal = {PVLDB},
        series = {{VLDB} '08},
        volume = {1},
        number = {1},
        pages = {1265--1276},
        doi = {10.14778/1454159.1454224},
        url = {https://doi.org/10.14778/1454159.1454224},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 89 citing papers.

Rank Citing Paper Year Venue Pagerank
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
51 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004291425
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031680027
295 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022238183
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
642 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015395331
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
710 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014715033
769 A Comparison of Join Algorithms for Log Processing in MapReduce 2010 SIGMOD 0.00014166872
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
872 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013486409
945 Fast Personalized PageRank on MapReduce 2011 SIGMOD 0.00013066956
954 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012997301
993 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012789598
1,021 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012606673
1,067 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012327784
1,108 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012145154
1,116 Large Graph Processing in the Cloud 2010 SIGMOD 0.0001210972
1,149 A Case for A Collaborative Query Management System 2009 CIDR 0.0001194728
1,264 SQL/MapReduce: A practical approach to self-describing, polymorphic, and parallelizable user-defined functions 2009 VLDB 0.00011416393
1,436 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010797443
1,442 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.00010778486
1,468 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010686496
1,621 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00010203114
1,694 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.9953156e-05
1,794 Distributed Data-Parallel Computing Using a High-Level Programming Language 2009 SIGMOD 9.7399162e-05
1,852 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.6134443e-05
1,925 POLARIS: The Distributed SQL Engine in Azure Synapse 2020 VLDB 9.4764398e-05
2,094 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 9.1819738e-05
2,136 Combining User Interaction, Speculative Query Execution and Sampling in the DICE System 2014 VLDB 9.1126926e-05
2,164 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 9.0521951e-05
2,196 Spinning Fast Iterative Data Flows 2012 VLDB 8.9704984e-05
2,265 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.8398946e-05
2,420 Lero: A Learning-to-Rank Query Optimizer 2023 VLDB 8.605257e-05
2,477 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.5239378e-05
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
2,567 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.4098241e-05
2,717 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.2102313e-05
2,822 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 8.0898536e-05
3,117 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7382559e-05
3,137 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.7204167e-05
3,466 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.3909785e-05
3,605 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2640711e-05
3,749 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.1539955e-05
3,964 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9855158e-05
3,998 Deploying a Steered Query Optimizer in Production at Microsoft 2022 SIGMOD 6.9676473e-05
4,244 Algorithmic Aspects of Parallel Query Processing 2018 SIGMOD 6.8062787e-05
4,389 Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka 2021 SIGMOD 6.7304854e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
72 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.00037695852
Previous Page 1 / 1 Next

Semantically Similar Papers