DBScholar

Back to papers

Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins

Summary: Process nested Parquet in relational engines via on-the-fly join-key generation to reconstruct nesting without materializing an internal format. Scans flat Parquet columns and uses on-the-fly joins to rebuild nesting, delivering vastly faster analytics. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
ha95fbe7c1d03ec31
Venue
SIGMOD
Year
2025
Pagerank
4.9769913e-05
Overall Rank
11,200 | 24.73%
DOI
10.1145/3725329
PDF
Download (CC BY 4.0)

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{rey_sigmod25,
        title = {{Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins}},
        author = {Rey, Alice and Rieger, Maximilian and Neumann, Thomas},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725329},
        url = {https://dl.acm.org/doi/10.1145/3725329},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004314366
71 DuckDB: an Embeddable Analytical Database 2019 SIGMOD 0.00037724477
86 XMark: A Benchmark for XML Data Management 2002 VLDB 0.00035793969
141 Relational Databases for Querying XML Documents: Limitations and Opportunities 1999 VLDB 0.0002928095
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024844328
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023136934
326 Storing Semistructured Data with STORED 1999 SIGMOD 0.00020953829
362 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00019999596
373 Umbra: A Disk-Based System with In-Memory Performance 2020 CIDR 0.00019705706
921 AsterixDB: A Scalable, Open Source BDMS 2014 VLDB 0.00013064043
1,135 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00011886308
1,487 Photon: A Fast Query Engine for Lakehouse Systems 2022 SIGMOD 0.00010516813
2,317 BtrBlocks: Efficient Columnar Compression for Data Lakes 2023 SIGMOD 8.6492924e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9703078e-05
3,810 An Empirical Evaluation of Columnar Storage Formats 2024 VLDB 7.0064171e-05
3,917 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.925328e-05
4,638 A Deep Dive into Common Open Formats for Analytical DBMSs 2023 VLDB 6.4901967e-05
5,934 The Flatter, the Better: Query Compilation Based on the Flattening Transformation 2015 SIGMOD 5.9392173e-05
6,395 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.7981019e-05
6,440 Scalable Querying of Nested Data 2021 VLDB 5.7823948e-05
6,520 Exploiting Common Patterns for Tree-Structured Data 2017 SIGMOD 5.7552712e-05
7,547 Storing and Querying Tree-Structured Records in Dremel 2014 VLDB 5.4962678e-05
8,849 Columnar Formats for Schemaless LSM-based Document Stores 2022 VLDB 5.2640252e-05
Previous Page 1 / 1 Next

Semantically Similar Papers