DBScholar

Back to papers

Selecting Subexpressions to Materialize at Datacenter Scale

Summary: Selects cross-job subexpressions to materialize at datacenter scale, formulating the problem as ILP/bipartite labeling. BIG SUBS, a distributed vertex-centric algorithm, handles tens of thousands of jobs and cuts machine-hours by up to 40% on SCOPE workloads. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h6542a7c977054b41
Venue
VLDB
Year
2018
Pagerank
9.7303647e-05
Overall Rank
1,747 | 88.26%
DOI
10.14778/3192965.3192971

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jindal_vldb18,
        title = {{Selecting Subexpressions to Materialize at Datacenter Scale}},
        author = {Jindal, Alekh and Karanasos, Konstantinos and Rao, Sriram and Patel, Hiren},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {7},
        pages = {800--812},
        doi = {10.14778/3192965.3192971},
        url = {https://doi.org/10.14778/3192965.3192971},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010208225
2,787 EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views 2022 SIGMOD 8.0121053e-05
2,833 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 7.9539771e-05
3,297 Automated Verification of Query Equivalence Using Satisfiability Modulo Theories 2019 VLDB 7.4421962e-05
3,321 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.427185e-05
3,680 openGauss: An Autonomous Database System 2021 VLDB 7.1016555e-05
4,007 Automated Generation of Materialized Views in Oracle 2020 VLDB 6.8561069e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
5,438 Eraser: Eliminating Performance Regression on Learned Query Optimizer 2024 VLDB 6.1278045e-05
5,451 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 6.1238308e-05
5,774 Optimizing Data Pipelines for Machine Learning in Feature Stores 2023 VLDB 5.9959266e-05
6,248 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.8345657e-05
6,758 Multi-Tenant Cloud Data Services: State-of-the-Art, Challenges and Opportunities 2022 SIGMOD 5.6877154e-05
6,801 Intermittent Query Processing 2019 VLDB 5.6776677e-05
6,820 CrocodileDB: Efficient Database Execution through Intelligent Deferment 2020 CIDR 5.6708756e-05
6,971 Jigsaw: A Data Storage and Query Processing Engine for Irregular Table Partitioning 2021 SIGMOD 5.628441e-05
7,160 Sibyl: Forecasting Time-Evolving Query Workloads 2024 SIGMOD 5.5945786e-05
7,344 Scalable Multi-Query Execution using Reinforcement Learning 2021 SIGMOD 5.5455974e-05
7,570 SIEVE: Effective Filtered Vector Search with Collection of Indexes 2025 VLDB 5.4929234e-05
7,604 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4846038e-05
8,349 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.3485189e-05
8,542 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.3188219e-05
8,704 View Selection over Knowledge Graphs in Triple Stores 2021 VLDB 5.2880532e-05
8,972 GEqO: ML-Accelerated Semantic Equivalence Detection 2023 SIGMOD 5.2469075e-05
9,028 Optimizing the cloud? Don't train models. Build oracles! 2024 CIDR 5.2330697e-05
9,098 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.2258409e-05
9,367 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.1845217e-05
9,532 UniView: A Unified Autonomous Materialized View Management System for Various Databases 2024 VLDB 5.1648692e-05
9,748 When sweet and cute isn't enough anymore: Solving scalability issues in Python Pandas with Grizzly 2020 CIDR 5.1325223e-05
9,999 SparkCruise: Handsfree Computation Reuse in Spark 2019 VLDB 5.0964027e-05
10,160 Generating Application-Specific Data Layouts for In-memory Databases 2019 VLDB 5.0683727e-05
10,905 Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting 2026 VLDB 4.9769913e-05
11,437 Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining 2025 VLDB 4.9769913e-05
11,467 Oligolithic Cross-task Optimizations across Isolated Workloads* 2024 CIDR 4.9769913e-05
11,720 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark 2023 SIGMOD 4.9769913e-05
11,854 Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications 2022 SIGMOD 4.9769913e-05
13,718 PikePlace: Generating Intelligence for Marketplace Datasets 2023 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012087459
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071056708
25 NiagaraCQ: A Scalable Continuous Query System for Internet Databases 2000 SIGMOD 0.00053906051
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049821554
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035340164
129 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.00030395767
511 TelegraphCQ: Continuous Dataflow Processing 2003 SIGMOD 0.00017058115
704 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014653079
762 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014134432
780 The case against specialized graph analytics engines 2015 CIDR 0.00014037973
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013642066
1,132 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011893781
1,549 Data Warehouse Configuration 1997 VLDB 0.00010287654
1,925 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3643089e-05
2,086 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0677901e-05
2,225 Shared Workload Optimization 2014 VLDB 8.8062552e-05
2,356 Vertexica: Your Relational Friend for Graph Analytics! 2014 VLDB 8.5852943e-05
2,448 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4508839e-05
3,215 Efficient and Provable Multi-Query Optimization 2017 PODS 7.5233633e-05
3,217 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5199872e-05
3,544 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2108612e-05
4,472 Recurring Job Optimization in Scope 2012 SIGMOD 6.5822904e-05
5,091 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 6.2775613e-05
5,836 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9732317e-05
7,365 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.5401198e-05
8,319 View Selection in Semantic Web Databases 2012 VLDB 5.3546235e-05
9,162 Delta: Scalable Data Dissemination under Capacity Constraints 2014 VLDB 5.2152285e-05
Previous Page 1 / 1 Next

Semantically Similar Papers