DBScholar

Back to papers

ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data

Summary: GPU-parallel delimiter parsing via FSM simulation eliminates the sequential context pass, supporting expressive formats such as quoted fields. Scales to thousands of cores at 14.2 GB/s, with streaming PCIe transfers enabling 4.8 GB end-to-end parsing in 0.44 s. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12449
Venue
VLDB
Year
2020
Pagerank
5.668373e-05
Overall Rank
7,217 | 50.49%
DOI
10.14778/3377369.3377372

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{stehle_vldb20,
        title = {{ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data}},
        author = {Stehle, Elias and Jacobsen, Hans-Arno},
        journal = {PVLDB},
        series = {{VLDB} '20},
        volume = {13},
        number = {5},
        pages = {616--628},
        doi = {10.14778/3377369.3377372},
        url = {https://doi.org/10.14778/3377369.3377372},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Rank Citing Paper Year Venue Pagerank
6,012 Parallel Index-based Stream Join on a Multicore CPU 2020 SIGMOD 6.0102675e-05
7,766 Scalable Structural Index Construction for JSON Analytics 2021 VLDB 5.549481e-05
9,436 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.2692207e-05
9,989 GpJSON: High-performance JSON Data Processing on GPUs 2025 VLDB 5.1826377e-05
10,102 Distributed Stream KNN Join 2021 SIGMOD 5.1488731e-05
10,761 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.0001232307
1,903 Instant Loading for Main Memory Databases 2013 VLDB 9.5049156e-05
1,927 Here are my Data Files. Here are my Queries. Where are my Results? 2011 CIDR 9.4703074e-05
2,389 Parallel Data Analysis Directly on Scientific File Formats 2014 SIGMOD 8.6439053e-05
2,391 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.6413407e-05
2,414 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.6078841e-05
2,435 Parallel In-Situ Data Processing with Speculative Loading 2014 SIGMOD 8.5811022e-05
2,667 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.2756346e-05
2,913 Constance: An Intelligent Data Lake System 2016 SIGMOD 7.9684737e-05
3,062 Adaptive Query Processing on RAW Data 2014 VLDB 7.8037446e-05
3,074 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 7.7844208e-05
3,378 A Hybrid B+-tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms 2016 SIGMOD 7.4587887e-05
3,638 Fast Queries Over Heterogeneous Data Through Engine Customization 2016 VLDB 7.2338361e-05
4,090 The bionic DBMS is coming, but what will it look like? 2013 CIDR 6.9096664e-05
4,764 ReCache: Reactive Caching for Fast Analytics over Heterogeneous Data 2018 VLDB 6.51896e-05
5,048 FAD.js: Fast JSON Data Access Using JIT-based Speculative Optimizations 2017 VLDB 6.3870952e-05
7,164 FishStore: Fast Ingestion and Indexing of Raw Data 2019 VLDB 5.6848225e-05
Previous Page 1 / 1 Next

Semantically Similar Papers