Back to papers
Filter Before You Parse: Faster Analytics on Raw Data with Sparser
Summary: Raw filtering applies predicates to the raw bytestream before parsing, dramatically reducing parsing overhead. SIMD RF cascades with a lightweight optimizer let Sparser pick the best cascade per data/format (JSON/Avro/Parquet), delivering up to 22x parser and 9x end-to-end speedups.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h0e237e9ceb13cfd7
Venue
VLDB
Year
2018
Pagerank
8.5972138e-05
Overall Rank
2,348 | 84.23%
DOI
10.14778/3236187.3236207
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@article{palkar_vldb18,
title = {{Filter Before You Parse: Faster Analytics on Raw Data with Sparser}},
author = {Palkar, Shoumik and Abuzaid, Firas and Bailis, Peter and Zaharia, Matei},
journal = {PVLDB},
series = {{VLDB} '18},
volume = {11},
number = {11},
pages = {1576--1589},
doi = {10.14778/3236187.3236207},
url = {https://doi.org/10.14778/3236187.3236207},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 18 of 18 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
1,669
SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle
2020
CIDR
9.9324573e-05
2,735
Speculative Distributed CSV Data Parsing for Big Data Analytics
2019
SIGMOD
8.0731466e-05
3,917
JSON Tiles: Fast Analytics on Semi-Structured Data
2021
SIGMOD
6.925328e-05
4,575
Accelerating Raw Data Analysis with the ACCORDA Software and Hardware Architecture
2019
VLDB
6.5234168e-05
4,596
AS-Parser: Log Parsing Based on Adaptive Segmentation
2023
SIGMOD
6.5074617e-05
5,959
Cheetah: Accelerating Database Queries with Switch Pruning
2020
SIGMOD
5.9310224e-05
6,064
ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data
2020
VLDB
5.8954903e-05
6,395
Selection Pushdown in Column Stores using Bit Manipulation Instructions
2023
SIGMOD
5.7981019e-05
7,931
Scalable Structural Index Construction for JSON Analytics
2021
VLDB
5.4230698e-05
8,093
Stackless Processing of Streamed Trees
2021
PODS
5.3917406e-05
8,961
FishStore: Faster Ingestion with Subset Hashing
2019
SIGMOD
5.2490978e-05
9,330
Dynamic Speculative Optimizations for SQL Compilation in Apache Spark
2020
VLDB
5.1918883e-05
9,545
GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example
2023
SIGMOD
5.1600923e-05
10,139
GpJSON: High-performance JSON Data Processing on GPUs
2025
VLDB
5.0725068e-05
10,312
Fast and Scalable Data Transfer Across Data Systems
2025
SIGMOD
5.0376863e-05
10,342
Zed: Leveraging Data Types to Process Eclectic Data
2023
CIDR
5.0230745e-05
10,944
OmniTable: A Unified Wide-Table System for Petabyte-Scale LLM Data Curation and Exploration
2026
VLDB
4.9769913e-05
11,714
dsJSON: A Distributed SQL JSON Processor
2023
SIGMOD
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
23
Spark SQL: Relational Data Processing in Spark
2015
SIGMOD
0.00055384955
827
Adaptive Ordering of Pipelined Stream Filters
2004
SIGMOD
0.00013629035
1,070
NoDB: Efficient Query Execution on Raw Data Files
2012
SIGMOD
0.00012179575
1,380
H2O: A Hands-free Adaptive Store
2014
SIGMOD
0.00010858313
1,745
Sinew: A SQL System for Multi-Structured Data
2014
SIGMOD
9.7324334e-05
1,927
Here are my Data Files. Here are my Queries. Where are my Results?
2011
CIDR
9.362697e-05
1,933
Instant Loading for Main Memory Databases
2013
VLDB
9.3430609e-05
2,331
Parallel In-Situ Data Processing with Speculative Loading
2014
SIGMOD
8.6253336e-05
2,437
Mison: A Fast JSON Parser for Data Analytics
2017
VLDB
8.4668658e-05
3,071
Adaptive Query Processing on RAW Data
2014
VLDB
7.6789108e-05
3,102
Micro Adaptivity in Vectorwise
2013
SIGMOD
7.6469535e-05
3,445
Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing
2017
VLDB
7.2911896e-05
3,505
Fast Queries Over Heterogeneous Data Through Engine Customization
2016
VLDB
7.2474175e-05
6,281
Just-In-Time Data Virtualization: Lightweight Data Management with ViDa
2015
CIDR
5.8236382e-05
7,924
AFilter: Adaptable XML Filtering with Prefix-Caching and Suffix-Clustering
2006
VLDB
5.4238547e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
9,330
Dynamic Speculative Optimizations for SQL Compilation in Apache Spark
2020
VLDB
2
10,146
A four-dimensional Analysis of Partitioned Approximate Filters
2021
VLDB
3
4,575
Accelerating Raw Data Analysis with the ACCORDA Software and Hardware Architecture
2019
VLDB
4
2,437
Mison: A Fast JSON Parser for Data Analytics
2017
VLDB
5
10,890
One Pass to Parse Them All: Fused Parallel CSV Processing
2026
VLDB
6
5,514
Runtime-Extensible Parsers
2025
CIDR
7
2,331
Parallel In-Situ Data Processing with Speculative Loading
2014
SIGMOD
8
3,071
Adaptive Query Processing on RAW Data
2014
VLDB
9
6,064
ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data
2020
VLDB
10
2,735
Speculative Distributed CSV Data Parsing for Big Data Analytics
2019
SIGMOD