DBScholar

Back to papers

Petabyte-Scale Row-Level Operations in Data Lakehouses

Summary: Adds petabyte-scale row-level updates/deletes to Iceberg+Spark via file-level materialization or lazy equality/position deletes. Avoids shuffles with storage-partitioned joins, cuts write amplification with runtime filters, and uses adaptive writes; shows use-case tradeoffs and ~10× gains. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h3debeaa935333da5
Venue
VLDB
Year
2024
Pagerank
5.6441822e-05
Overall Rank
6,920 | 53.48%
DOI
10.14778/3685800.3685834

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{okolnychyi_vldb24,
        title = {{Petabyte-Scale Row-Level Operations in Data Lakehouses}},
        author = {Okolnychyi, Anton and Sun, Chao and Tanimura, Kazuyuki and Spitzer, Russell and Blue, Ryan and Ho, Szehon and Gu, Yufei and Lakkundi, Vishwanath and Tsai, DB},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {12},
        pages = {4159--4172},
        doi = {10.14778/3685800.3685834},
        url = {https://doi.org/10.14778/3685800.3685834},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 21 of 21 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
12 C-Store: A Column-oriented DBMS 2005 VLDB 0.00068998927
14 MonetDB/X100: Hyper-Pipelining Query Execution 2005 CIDR 0.00064031282
19 A Critique of ANSI SQL Isolation Levels 1995 SIGMOD 0.00058781151
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
52 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00041219077
59 Differential Files: Their Application To The Maintenance Of Large Data Bases 1976 SIGMOD 0.00039605949
216 Small Materialized Aggregates: A Light Weight Index Structure for Data Warehousing 1998 VLDB 0.00024485024
233 Amazon Redshift and the Case for Simpler Data Warehouses 2015 SIGMOD 0.00023783585
237 Serializable Isolation for Snapshot Databases 2008 SIGMOD 0.00023660039
459 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017856221
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
950 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012895553
1,135 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00011891907
1,632 Positional Update Handling in Column Stores 2010 SIGMOD 0.00010021251
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
3,446 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing 2019 SIGMOD 7.2942885e-05
3,928 Big Metadata: When Metadata is Big Data 2021 VLDB 6.9192051e-05
4,332 Analyzing and Comparing Lakehouse Storage Systems 2023 CIDR 6.6587744e-05
5,504 Magnet: Push-based Shuffle Service for Large-scale Data Processing 2020 VLDB 6.1018427e-05
5,817 VectorH: Taking SQL-on-Hadoop to the Next Level 2016 SIGMOD 5.9840356e-05
8,208 LST-Bench: Benchmarking Log-Structured Tables in the Cloud 2024 SIGMOD 5.3775598e-05
Previous Page 1 / 1 Next

Semantically Similar Papers