DBScholar

Back to papers

Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins

Summary: Process nested Parquet in relational engines via on-the-fly join-key generation to reconstruct nesting without materializing an internal format. Scans flat Parquet columns and uses on-the-fly joins to rebuild nesting, delivering vastly faster analytics. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
ha95fbe7c1d03ec31
Venue
SIGMOD
Year
2025
Pagerank
4.9793485e-05
Overall Rank
11,191 | 24.76%
DOI
10.1145/3725329

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{rey_sigmod25,
        title = {{Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins}},
        author = {Rey, Alice and Rieger, Maximilian and Neumann, Thomas},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725329},
        url = {https://dl.acm.org/doi/10.1145/3725329},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
71 DuckDB: an Embeddable Analytical Database 2019 SIGMOD 0.00037720227
86 XMark: A Benchmark for XML Data Management 2002 VLDB 0.00035810622
141 Relational Databases for Querying XML Documents: Limitations and Opportunities 1999 VLDB 0.00029294114
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
326 Storing Semistructured Data with STORED 1999 SIGMOD 0.00020963333
361 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020006406
373 Umbra: A Disk-Based System with In-Memory Performance 2020 CIDR 0.00019711632
922 AsterixDB: A Scalable, Open Source BDMS 2014 VLDB 0.00013068048
1,135 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00011891907
1,487 Photon: A Fast Query Engine for Lakehouse Systems 2022 SIGMOD 0.00010521722
2,314 BtrBlocks: Efficient Columnar Compression for Data Lakes 2023 SIGMOD 8.6533171e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
3,808 An Empirical Evaluation of Columnar Storage Formats 2024 VLDB 7.0097354e-05
3,916 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.9285787e-05
4,635 A Deep Dive into Common Open Formats for Analytical DBMSs 2023 VLDB 6.4932705e-05
5,934 The Flatter, the Better: Query Compilation Based on the Flattening Transformation 2015 SIGMOD 5.9420302e-05
6,392 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.800848e-05
6,438 Scalable Querying of Nested Data 2021 VLDB 5.7851309e-05
6,518 Exploiting Common Patterns for Tree-Structured Data 2017 SIGMOD 5.757997e-05
7,540 Storing and Querying Tree-Structured Records in Dremel 2014 VLDB 5.4988709e-05
8,840 Columnar Formats for Schemaless LSM-based Document Stores 2022 VLDB 5.2665183e-05
Previous Page 1 / 1 Next

Semantically Similar Papers