DBScholar

Back to papers

Selecting Subexpressions to Materialize at Datacenter Scale

Summary: Selects cross-job subexpressions to materialize at datacenter scale, formulating the problem as ILP/bipartite labeling. BIG SUBS, a distributed vertex-centric algorithm, handles tens of thousands of jobs and cuts machine-hours by up to 40% on SCOPE workloads. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h6542a7c977054b41
Venue
VLDB
Year
2018
Pagerank
9.7343818e-05
Overall Rank
1,745 | 88.27%
DOI
10.14778/3192965.3192971

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jindal_vldb18,
        title = {{Selecting Subexpressions to Materialize at Datacenter Scale}},
        author = {Jindal, Alekh and Karanasos, Konstantinos and Rao, Sriram and Patel, Hiren},
        journal = {PVLDB},
        series = {{VLDB} '18},
        volume = {11},
        number = {7},
        pages = {800--812},
        doi = {10.14778/3192965.3192971},
        url = {https://doi.org/10.14778/3192965.3192971},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
2,787 EVA: A Symbolic Approach to Accelerating Exploratory Video Analytics with Materialized Views 2022 SIGMOD 8.0158999e-05
2,834 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings 2020 SIGMOD 7.9560627e-05
3,297 Automated Verification of Query Equivalence Using Satisfiability Modulo Theories 2019 VLDB 7.4452842e-05
3,336 GenRewrite: Query Rewriting via Large Language Models 2026 SIGMOD 7.4137763e-05
3,682 openGauss: An Autonomous Database System 2021 VLDB 7.1013922e-05
4,007 Automated Generation of Materialized Views in Oracle 2020 VLDB 6.8592987e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
5,446 Crystal: A Unified Cache Storage System for Analytical Databases 2021 VLDB 6.1266228e-05
5,456 Eraser: Eliminating Performance Regression on Learned Query Optimizer 2024 VLDB 6.1239873e-05
5,773 Optimizing Data Pipelines for Machine Learning in Feature Stores 2023 VLDB 5.9987664e-05
6,245 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.837329e-05
6,753 Multi-Tenant Cloud Data Services: State-of-the-Art, Challenges and Opportunities 2022 SIGMOD 5.6904085e-05
6,796 Intermittent Query Processing 2019 VLDB 5.6803529e-05
6,814 CrocodileDB: Efficient Database Execution through Intelligent Deferment 2020 CIDR 5.6735614e-05
6,970 Jigsaw: A Data Storage and Query Processing Engine for Irregular Table Partitioning 2021 SIGMOD 5.6311067e-05
7,157 Sibyl: Forecasting Time-Evolving Query Workloads 2024 SIGMOD 5.5972283e-05
7,341 Scalable Multi-Query Execution using Reinforcement Learning 2021 SIGMOD 5.5481233e-05
7,564 SIEVE: Effective Filtered Vector Search with Collection of Indexes 2025 VLDB 5.4955249e-05
7,598 SageDB: An Instance-Optimized Data Analytics System 2022 VLDB 5.4871733e-05
8,346 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.3510511e-05
8,535 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.320973e-05
8,696 View Selection over Knowledge Graphs in Triple Stores 2021 VLDB 5.2905577e-05
8,962 GEqO: ML-Accelerated Semantic Equivalence Detection 2023 SIGMOD 5.2493925e-05
9,020 Optimizing the cloud? Don't train models. Build oracles! 2024 CIDR 5.2355482e-05
9,088 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.2283159e-05
9,358 Phoebe: A Learning-based Checkpoint Optimizer 2021 VLDB 5.1869771e-05
9,520 UniView: A Unified Autonomous Materialized View Management System for Various Databases 2024 VLDB 5.1673153e-05
9,743 When sweet and cute isn't enough anymore: Solving scalability issues in Python Pandas with Grizzly 2020 CIDR 5.1349531e-05
9,993 SparkCruise: Handsfree Computation Reuse in Spark 2019 VLDB 5.0988164e-05
10,156 Generating Application-Specific Data Layouts for In-memory Databases 2019 VLDB 5.0707546e-05
10,896 Benchmarking the Full Pipeline of Materialized-View-Based Query Rewriting 2026 VLDB 4.9793485e-05
11,431 Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining 2025 VLDB 4.9793485e-05
11,461 Oligolithic Cross-task Optimizations across Isolated Workloads* 2024 CIDR 4.9793485e-05
11,714 QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark 2023 SIGMOD 4.9793485e-05
11,848 Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications 2022 SIGMOD 4.9793485e-05
13,713 PikePlace: Generating Intelligence for Marketplace Datasets 2023 VLDB -
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012092602
11 Implementing Data Cubes Efficiently 1996 SIGMOD 0.00071084324
25 NiagaraCQ: A Scalable Continuous Query System for Internet Databases 2000 SIGMOD 0.00053930011
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035351639
129 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.0003040756
510 TelegraphCQ: Continuous Dataflow Processing 2003 SIGMOD 0.00017065714
705 Main-Memory Scan Sharing For Multi-Core CPUs 2008 VLDB 0.00014657491
762 Cooperative Scans: Dynamic Bandwidth Sharing in a DBMS 2007 VLDB 0.00014140446
778 The case against specialized graph analytics engines 2015 CIDR 0.00014043807
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
1,131 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.0001189909
1,547 Data Warehouse Configuration 1997 VLDB 0.00010292395
1,924 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3687009e-05
2,084 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0716512e-05
2,225 Shared Workload Optimization 2014 VLDB 8.8081001e-05
2,355 Vertexica: Your Relational Friend for Graph Analytics! 2014 VLDB 8.5893186e-05
2,446 Azure Data Lake Store: A Hyperscale Distributed File Service for Big Data Analytics 2017 SIGMOD 8.4547121e-05
3,214 Efficient and Provable Multi-Query Optimization 2017 PODS 7.5269127e-05
3,216 Multi-Query Optimization in MapReduce Framework 2014 VLDB 7.5234702e-05
3,545 Computation Reuse in Analytics Job Service at Microsoft 2018 SIGMOD 7.2134803e-05
4,470 Recurring Job Optimization in Scope 2012 SIGMOD 6.5853561e-05
5,088 ROBUS: Fair Cache Allocation for Data-parallel Workloads 2017 SIGMOD 6.2805079e-05
5,834 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9759267e-05
7,360 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.5427171e-05
8,313 View Selection in Semantic Web Databases 2012 VLDB 5.3571595e-05
9,153 Delta: Scalable Data Dissemination under Capacity Constraints 2014 VLDB 5.2176985e-05
Previous Page 1 / 1 Next

Semantically Similar Papers