DBScholar

Back to papers

SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets

Summary: SCOPE is a declarative, extensible scripting language for data analysis on clusters, hiding parallelism from users. It provides SQL-like modeling with joins and aggregates, plus user-defined operators (extractors, processors, reducers, combiners), nesting, and stepwise plans compiled into parallel execution. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9943
Venue
VLDB
Year
2008
Pagerank
0.00051174276
Overall Rank
30 | 99.80%
DOI
10.14778/1454159.1454224

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chaiken_vldb08,
        title = {{SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets}},
        author = {Chaiken, Ronnie and Jenkins, Bob and Larson, Per-Åke and Ramsey, Bill and Shakib, Darren and Weaver, Simon and Zhou, Jingren},
        journal = {PVLDB},
        series = {{VLDB} '08},
        volume = {1},
        number = {1},
        pages = {1265--1276},
        doi = {10.14778/1454159.1454224},
        url = {https://doi.org/10.14778/1454159.1454224},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 39 of 89 citing papers.

Rank Citing Paper Year Venue Pagerank
4,400 Recurring Job Optimization in Scope 2012 SIGMOD 6.7239302e-05
4,412 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.7168614e-05
4,696 Error-bounded Sampling for Analytics on Big Sparse Data 2014 VLDB 6.557612e-05
4,918 Massive Scale-out of Expensive Continuous Queries 2011 VLDB 6.4469122e-05
5,059 Steering Query Optimizers: A Practical Take on Big Data Workloads 2021 SIGMOD 6.3807509e-05
5,182 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.3287692e-05
5,257 Efficient Estimation of Inclusion Coefficient using HyperLogLog Sketches 2018 VLDB 6.2971456e-05
5,279 Mining Document Collections to Facilitate Accurate Approximate Entity Matching 2009 VLDB 6.2849137e-05
5,344 Discovering Related Data At Scale 2021 VLDB 6.2591104e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
6,045 Helios: Hyperscale Indexing for the Cloud & Edge 2020 VLDB 5.9957544e-05
6,102 AutoExecutor: Predictive Parallelism for Spark SQL Queries 2021 VLDB 5.976708e-05
6,377 Scalable Querying of Nested Data 2021 VLDB 5.8931544e-05
6,462 PilotScope: Steering Databases with Machine Learning Drivers 2024 VLDB 5.8717744e-05
6,946 KEA: Tuning an Exabyte-Scale Data Infrastructure 2021 SIGMOD 5.7309848e-05
7,176 Interactive Demonstration of Probabilistic Predicates 2018 SIGMOD 5.6816132e-05
7,182 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.6793679e-05
7,551 Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams 2022 VLDB 5.6006414e-05
7,619 AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft 2020 VLDB 5.5810604e-05
7,690 Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes 2021 SIGMOD 5.5669307e-05
7,762 Runtime Variation in Big Data Analytics 2023 SIGMOD 5.5501898e-05
7,878 Dependency-Driven Analytics: a Compass for Uncharted Data Oceans 2017 CIDR 5.5246167e-05
8,001 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.508791e-05
8,040 PerfGuard: Deploying ML-for-Systems without Performance Regressions, Almost! 2021 VLDB 5.5018396e-05
8,097 Couchbase Analytics: NoETL for Scalable NoSQL Data Analysis 2019 VLDB 5.4875142e-05
8,108 Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters 2019 VLDB 5.4850569e-05
8,175 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.4737932e-05
8,309 Spur: Mitigating Slow Instances in Large-Scale Streaming Pipelines 2020 SIGMOD 5.4560857e-05
8,448 Towards Query Optimizer as a Service (QOaaS) in a Unified LakeHouse Ecosystem: Can One QO Rule Them All? 2025 CIDR 5.4235725e-05
9,224 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.3035811e-05
9,274 DataGarage: Warehousing Massive Performance Data on Commodity Servers 2010 VLDB 5.2941406e-05
9,478 Parallel Query Processing: To Separate Communication from Computation 2022 SIGMOD 5.2634238e-05
9,670 Optimistic Recovery for Iterative Dataflows in Action 2015 SIGMOD 5.2389261e-05
10,690 Dynamic Pruning for Recursive Joins 2025 SIGMOD 5.093636e-05
11,151 Flux: Decoupled Auto-Scaling for Heterogeneous Query Workload in Alibaba AnalyticDB 2024 SIGMOD 5.093636e-05
11,728 Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters 2021 VLDB 5.093636e-05
11,885 Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology 2019 VLDB 5.093636e-05
12,306 Declarative Error Management for Robust Data-Intensive Applications 2012 SIGMOD 5.093636e-05
12,421 Indexing Multi-dimensional Data in a Cloud System 2010 SIGMOD 5.093636e-05
Previous Page 2 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010686205
72 Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters 2007 SIGMOD 0.00037695852
Previous Page 1 / 1 Next

Semantically Similar Papers