DBScholar

Back to papers

Selecting Subexpressions to Materialize at Datacenter Scale

Summary: Selects cross-job subexpressions to materialize at datacenter scale, formulating the problem as ILP/bipartite labeling. BIG SUBS, a distributed vertex-centric algorithm, handles tens of thousands of jobs and cuts machine-hours by up to 40% on SCOPE workloads. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
11973
Venue
VLDB
Year
2018
Pagerank
9.8079546e-05
Overall Rank
1,765 | 87.90%
DOI
10.14778/3192965.3192971

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jindal_vldb18,
        title = {{Selecting Subexpressions to Materialize at Datacenter Scale}},
        author = {Jindal, Alekh and Karanasos, Konstantinos and Rao, Sriram and Patel, Hiren},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {7},
        pages = {800--812},
        doi = {10.14778/3192965.3192971},
        url = {https://doi.org/10.14778/3192965.3192971},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 36 of 36 citing papers.

Rank Citing Paper Year Venue Pagerank
1,569 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010335423
2,822 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 8.0898536e-05
2,933 EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views 2022 SIGMOD 7.9474026e-05
3,327 Automated Verification of Query Equivalence Using Satisfiability Modulo Theories 2019 VLDB 7.518491e-05
3,662 openGauss: An Autonomous Database System 2021 VLDB 7.2166682e-05
4,101 Automated Generation of Materialized Views in Oracle 2020 VLDB 6.9009734e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,363 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 6.7423909e-05
5,348 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 6.2562688e-05
5,573 Eraser: Eliminating Performance Regression on Learned Query Optimizer 2024 VLDB 6.1682747e-05
5,704 Optimizing Data Pipelines for Machine Learning in Feature Stores 2023 VLDB 6.1146371e-05
6,121 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.9688569e-05
6,684 CrocodileDB: Efficient Database Execution through Intelligent Deferment 2020 CIDR 5.8036476e-05
6,848 Jigsaw: A Data Storage and Query Processing Engine for Irregular Table Partitioning 2021 SIGMOD 5.7550624e-05
6,939 Multi-Tenant Cloud Data Services: State-of-the-Art, Challenges and Opportunities 2022 SIGMOD 5.7338637e-05
7,200 Intermittent Query Processing 2019 VLDB 5.6756294e-05
7,213 Scalable Multi-Query Execution using Reinforcement Learning 2021 SIGMOD 5.670422e-05
7,580 Sibyl: Forecasting Time-Evolving Query Workloads 2024 SIGMOD 5.5925285e-05
8,175 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.4737932e-05
8,323 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4539294e-05
8,439 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.4248071e-05
8,527 View Selection over Knowledge Graphs in Triple Stores 2021 VLDB 5.4119882e-05
8,798 GEqO: ML-Accelerated Semantic Equivalence Detection 2023 SIGMOD 5.3698781e-05
8,861 Optimizing the cloud? Don't train models. Build oracles! 2024 CIDR 5.355716e-05
8,927 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.3483178e-05
9,145 SIEVE: Effective Filtered Vector Search with Collection of Indexes 2025 VLDB 5.3150984e-05
9,224 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.3035811e-05
9,567 When sweet and cute isn't enough anymore: Solving scalability issues in Python Pandas with Grizzly 2020 CIDR 5.2528121e-05
9,854 SparkCruise: Handsfree Computation Reuse in Spark 2019 VLDB 5.2091816e-05
9,949 UniView: A Unified Autonomous Materialized View Management System for Various Databases 2024 VLDB 5.1915905e-05
9,966 Generating Application-Specific Data Layouts for In-memory Databases 2019 VLDB 5.1870939e-05
11,074 Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining 2025 VLDB 5.093636e-05
11,113 Oligolithic Cross-task Optimizations across Isolated Workloads* 2024 CIDR 5.093636e-05
11,399 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark 2023 SIGMOD 5.093636e-05
11,539 Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications 2022 SIGMOD 5.093636e-05
13,399 PikePlace: Generating Intelligence for Marketplace Datasets 2023 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012250108
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071822821
25 NiagaraCQ: A Scalable Continuous Query System for Internet Databases 2000 SIGMOD 0.00054667018
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
87 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035281619
128 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.0003072825
505 TelegraphCQ: Continuous Dataflow Processing 2003 SIGMOD 0.00017285498
696 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014891322
759 The case against specialized graph analytics engines 2015 CIDR 0.00014273591
761 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014254351
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
1,154 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011934202
1,543 Data Warehouse Configuration 1997 VLDB 0.00010413384
1,883 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.5421713e-05
2,151 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0784444e-05
2,276 Shared Workload Optimization 2014 VLDB 8.8196376e-05
2,433 Vertexica: Your Relational Friend for Graph Analytics! 2014 VLDB 8.5869161e-05
2,477 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.5239378e-05
3,163 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.6784171e-05
3,268 Efficient and Provable Multi-Query Optimization 2017 PODS 7.5810998e-05
3,605 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2640711e-05
4,400 Recurring Job Optimization in Scope 2012 SIGMOD 6.7239302e-05
5,720 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 6.1092082e-05
7,238 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.664642e-05
7,673 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 5.5705944e-05
8,143 View Selection in Semantic Web Databases 2012 VLDB 5.4801187e-05
8,992 Delta: Scalable Data Dissemination under Capacity Constraints 2014 VLDB 5.3373227e-05
Previous Page 1 / 1 Next

Semantically Similar Papers