DBScholar

Back to papers

Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins

Summary: Process nested Parquet in relational engines via on-the-fly join-key generation to reconstruct nesting without materializing an internal format. Scans flat Parquet columns and uses on-the-fly joins to rebuild nesting, delivering vastly faster analytics. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
7310
Venue
SIGMOD
Year
2025
Pagerank
5.093636e-05
Overall Rank
10,771 | 26.11%
DOI
10.1145/3725329

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{rey_sigmod25,
        title = {{Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and Joins}},
        author = {Rey, Alice and Rieger, Maximilian and Neumann, Thomas},
        series = {{SIGMOD} '25},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3725329},
        url = {https://dl.acm.org/doi/10.1145/3725329},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
51 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004291425
79 XMark: A Benchmark for XML Data Management 2002 VLDB 0.00036555628
103 DuckDB: an Embeddable Analytical Database 2019 SIGMOD 0.00034161428
137 Relational Databases for Querying XML Documents: Limitations and Opportunities 1999 VLDB 0.00029877792
209 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024932174
252 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023242719
318 Storing Semistructured Data with STORED 1999 SIGMOD 0.00021380325
360 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020182846
422 Umbra: A Disk-Based System with In-Memory Performance 2020 CIDR 0.00018732744
1,015 AsterixDB: A Scalable, Open Source BDMS 2014 VLDB 0.00012647763
1,453 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00010742227
1,824 Photon: A Fast Query Engine for Lakehouse Systems 2022 SIGMOD 9.6734544e-05
2,788 BtrBlocks: Efficient Columnar Compression for Data Lakes 2023 SIGMOD 8.1205155e-05
2,962 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9170451e-05
3,877 An Empirical Evaluation of Columnar Storage Formats 2024 VLDB 7.0532293e-05
4,069 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.9276175e-05
5,037 A Deep Dive into Common Open Formats for Analytical DBMSs 2023 VLDB 6.3914026e-05
5,814 The Flatter, the Better: Query Compilation Based on the Flattening Transformation 2015 SIGMOD 6.0779807e-05
6,377 Scalable Querying of Nested Data 2021 VLDB 5.8931544e-05
6,385 Exploiting Common Patterns for Tree-Structured Data 2017 SIGMOD 5.8899875e-05
7,139 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.6932481e-05
7,405 Storing and Querying Tree-Structured Records in Dremel 2014 VLDB 5.624456e-05
8,682 Columnar Formats for Schemaless LSM-based Document Stores 2022 VLDB 5.3857524e-05
Previous Page 1 / 1 Next

Semantically Similar Papers