DBScholar

Back to papers

Partition, Don’t Sort! Compression Boosters for Cloud Data Ingestion Pipelines

Summary: Rather than expensive global sorting, cluster similarly structured nested Dremel-encoded records at ingestion to create compressible partitions. A decision-tree–inspired clustering is up to 17.44× faster than partition-then-sort and yields up to 2× compression, while per-bucket sorting matches increasing-cardinality compression at lower ingestion cost. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13744
Venue
VLDB
Year
2024
Pagerank
5.093636e-05
Overall Rank
11,274 | 22.66%
DOI
10.14778/3681954.3682013

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{hansert_vldb24,
        title = {{Partition, Don’t Sort! Compression Boosters for Cloud Data Ingestion Pipelines}},
        author = {Hansert, Patrick and Michel, Sebastian},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {11},
        pages = {3456--3469},
        doi = {10.14778/3681954.3682013},
        url = {https://doi.org/10.14778/3681954.3682013},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
21 Similarity Search in High Dimensions via Hashing 1999 VLDB 0.00056760516
51 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004291425
60 Integrating Compression and Execution in Column-Oriented Database Systems 2006 SIGMOD 0.0003955489
259 Database Cracking 2007 CIDR 0.00023119313
354 Linear Clustering of Objects with Multiple Attributes 1990 SIGMOD 0.00020355578
476 The Making of TPC-DS 2006 VLDB 0.00017860667
520 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017136828
970 Sybase IQ Multiplex – Designed For Analytics 2004 VLDB 0.00012882125
1,453 Dremel: A Decade of Interactive SQL Analysis at Web Scale 2020 VLDB 0.00010742227
3,035 Instance-Optimized Data Layouts for Cloud Analytics Workloads 2021 SIGMOD 7.8297746e-05
3,106 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.7515666e-05
3,722 NET-FLi: On-the-fly Compression, Archiving and Indexing of Streaming Network Traffic 2010 VLDB 7.1726862e-05
4,069 JSON Tiles: Fast Analytics on Semi-Structured Data 2021 SIGMOD 6.9276175e-05
5,985 Column Partition and Permutation for Run Length Encoding in Columnar Databases 2020 SIGMOD 6.0185035e-05
6,042 Pando: Enhanced Data Skipping with Logical Data Partitioning 2023 VLDB 5.9970052e-05
6,385 Exploiting Common Patterns for Tree-Structured Data 2017 SIGMOD 5.8899875e-05
6,480 Proteus: Autonomous Adaptive Storage for Mixed Workloads 2022 SIGMOD 5.8669819e-05
6,640 Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters 2013 VLDB 5.8162526e-05
6,848 Jigsaw: A Data Storage and Query Processing Engine for Irregular Table Partitioning 2021 SIGMOD 5.7550624e-05
6,882 Rearranging Data to Maximize the Efficiency of Compression 1986 PODS 5.7469443e-05
6,896 Wide Table Layout Optimization based on Column Ordering and Duplication 2017 SIGMOD 5.7440611e-05
7,327 Reducing Ambiguity in Json Schema Discovery 2021 SIGMOD 5.6435519e-05
7,465 Automated Multidimensional Data Layouts in Amazon Redshift 2024 SIGMOD 5.6108826e-05
Previous Page 1 / 1 Next

Semantically Similar Papers