DBScholar

Back to papers

SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets

Summary: SCOPE is a declarative, extensible scripting language for data analysis on clusters, hiding parallelism from users. It provides SQL-like modeling with joins and aggregates, plus user-defined operators (extractors, processors, reducers, combiners), nesting, and stepwise plans compiled into parallel execution. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hdc248f8ae945036c
Venue
VLDB
Year
2008
Pagerank
0.00050495102
Overall Rank
30 | 99.81%
DOI
10.14778/1454159.1454224

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chaiken_vldb08,
        title = {{SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets}},
        author = {Chaiken, Ronnie and Jenkins, Bob and Larson, Per-Åke and Ramsey, Bill and Shakib, Darren and Weaver, Simon and Zhou, Jingren},
        journal = {PVLDB},
        series = {{VLDB} '08},
        volume = {1},
        number = {1},
        pages = {1265--1276},
        doi = {10.14778/1454159.1454224},
        url = {https://doi.org/10.14778/1454159.1454224},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 89 citing papers.

Rank Citing Paper Year Venue Pagerank
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
281 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022295232
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
651 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015130782
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
685 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014782777
799 A Comparison of Join Algorithms for Log Processing in MapReduce 2010 SIGMOD 0.00013889081
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
962 Fast Personalized PageRank on MapReduce 2011 SIGMOD 0.00012820478
977 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012731794
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012539827
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012376819
1,080 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012131978
1,082 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012122749
1,137 Large Graph Processing in the Cloud 2010 SIGMOD 0.00011870544
1,164 A Case for A Collaborative Query Management System 2009 CIDR 0.00011745383
1,276 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00011239266
1,285 SQL/MapReduce: A practical approach to self-describing, polymorphic, and parallelizable user-defined functions 2009 VLDB 0.00011201377
1,433 Towards a Learning Optimizer for Shared Clouds 2019 VLDB 0.00010677711
1,444 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.0001063181
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010569837
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8224587e-05
1,790 POLARIS: The Distributed SQL Engine in Azure Synapse 2020 VLDB 9.6272006e-05
1,810 Distributed Data-Parallel Computing Using a High-Level Programming Language 2009 SIGMOD 9.5892079e-05
1,883 Automated Partitioning Design in Parallel Database Systems 2011 SIGMOD 9.4391795e-05
2,123 Tupleware: "Big" Data, Big Analytics, Small Clusters 2015 CIDR 9.0060385e-05
2,178 Combining User Interaction, Speculative Query Execution and Sampling in the DICE System 2014 VLDB 8.9152266e-05
2,195 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8781177e-05
2,210 Lero: A Learning-to-Rank Query Optimizer 2023 VLDB 8.8257742e-05
2,227 Spinning Fast Iterative Data Flows 2012 VLDB 8.8021772e-05
2,282 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6964781e-05
2,446 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4547121e-05
2,573 Minimal MapReduce Algorithms 2013 SIGMOD 8.2821647e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2738285e-05
2,754 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0534972e-05
2,834 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 7.9560627e-05
3,033 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7337998e-05
3,073 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.6777283e-05
3,440 Advanced Partitioning Techniques for Massively Distributed Computation 2012 SIGMOD 7.2986853e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
3,823 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.0029055e-05
3,895 Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE 2019 VLDB 6.9393157e-05
3,949 Deploying a Steered Query Optimizer in Production at Microsoft 2022 SIGMOD 6.9052796e-05
4,337 Algorithmic Aspects of Parallel Query Processing 2018 SIGMOD 6.6541797e-05
4,424 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.6048065e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
75 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.0003704106
Previous Page 1 / 1 Next

Semantically Similar Papers