Speculative Distributed CSV Data Parsing for Big Data Analytics
Summary: Speculative distributed CSV parsing aligns field/record boundaries across chunks without context to enable parallel parsing. Robust syntax-error detection; Spark tests on 11k real-world datasets show substantial performance gains over prior parsers. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chang Ge (University of Waterloo)
- 2. Yinan Li (Microsoft)
- 3. Eric Eilebrecht (Microsoft)
- 4. Badrish Chandramouli (Microsoft)
- 5. Donald Kossmann (Microsoft)
BibTeX Citation
@inproceedings{ge_sigmod19,
title = {{Speculative Distributed CSV Data Parsing for Big Data Analytics}},
author = {Ge, Chang and Li, Yinan and Eilebrecht, Eric and Chandramouli, Badrish and Kossmann, Donald},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3319898},
url = {https://dl.acm.org/doi/10.1145/3299869.3319898},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 13 of 13 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 51 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.0004291425 |
| 1,070 | NoDB: Efficient Query Execution on Raw Data Files | 2012 | SIGMOD | 0.0001232307 |
| 1,903 | Instant Loading for Main Memory Databases | 2013 | VLDB | 9.5049156e-05 |
| 1,927 | Here are my Data Files. Here are my Queries. Where are my Results? | 2011 | CIDR | 9.4703074e-05 |
| 2,389 | Parallel Data Analysis Directly on Scientific File Formats | 2014 | SIGMOD | 8.6439053e-05 |
| 2,391 | Mison: A Fast JSON Parser for Data Analytics | 2017 | VLDB | 8.6413407e-05 |
| 2,414 | Filter Before You Parse: Faster Analytics on Raw Data with Sparser | 2018 | VLDB | 8.6078841e-05 |
| 2,435 | Parallel In-Situ Data Processing with Speculative Loading | 2014 | SIGMOD | 8.5811022e-05 |
| 3,062 | Adaptive Query Processing on RAW Data | 2014 | VLDB | 7.8037446e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,738 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB |
| 2 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 3 | 2,435 | Parallel In-Situ Data Processing with Speculative Loading | 2014 | SIGMOD |
| 4 | 5,535 | Runtime-Extensible Parsers | 2025 | CIDR |
| 5 | 11,625 | Accelerating Complex Analytics using Speculation | 2021 | CIDR |
| 6 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 7 | 7,015 | A framework for annotating CSV-like data | 2016 | VLDB |
| 8 | 9,167 | Dynamic Speculative Optimizations for SQL Compilation in Apache Spark | 2020 | VLDB |
| 9 | 7,217 | ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data | 2020 | VLDB |
| 10 | 2,414 | Filter Before You Parse: Faster Analytics on Raw Data with Sparser | 2018 | VLDB |