Parallel In-Situ Data Processing with Speculative Loading
Summary: SCANRAW is a parallel in-situ operator for raw files, merging loading with external tables to preserve zero time-to-query. It employs a parallel super-scalar pipeline with speculative loading to overlap queries and conversion, balancing CPU and I/O. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yu Cheng (University of California Merced)
- 2. Florin Rusu (University of California Merced)
BibTeX Citation
@inproceedings{cheng_sigmod14,
title = {{Parallel In-Situ Data Processing with Speculative Loading}},
author = {Cheng, Yu and Rusu, Florin},
series = {{SIGMOD} '14},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2588555.2593673},
url = {https://dl.acm.org/doi/10.1145/2588555.2593673},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 18 of 18 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 53 | Eddies: Continuously Adaptive Query Processing | 2000 | SIGMOD | 0.00041071971 |
| 868 | Requirements for Science Data Bases and SciDB | 2009 | CIDR | 0.00013504796 |
| 1,024 | The DataPath System: A Data-Centric Analytic Processing Engine for Large Data Warehouses | 2010 | SIGMOD | 0.0001258839 |
| 1,070 | NoDB: Efficient Query Execution on Raw Data Files | 2012 | SIGMOD | 0.0001232307 |
| 1,811 | Merging What's Cracked, Cracking What's Merged: Adaptive Indexing in Main-Memory Column-Stores | 2011 | VLDB | 9.698026e-05 |
| 1,903 | Instant Loading for Main Memory Databases | 2013 | VLDB | 9.5049156e-05 |
| 1,924 | The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds | 2011 | VLDB | 9.4775501e-05 |
| 1,927 | Here are my Data Files. Here are my Queries. Where are my Results? | 2011 | CIDR | 9.4703074e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,337 | Multi-dimensional Resource Scheduling for Parallel Queries | 1996 | SIGMOD |
| 2 | 1,927 | Here are my Data Files. Here are my Queries. Where are my Results? | 2011 | CIDR |
| 3 | 12,461 | Ten Thousand SQLs: Parallel Keyword Queries Computing | 2010 | VLDB |
| 4 | 6,255 | Elastic Pipelining in an In-Memory Database Cluster | 2016 | SIGMOD |
| 5 | 3,413 | Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing | 2017 | VLDB |
| 6 | 7,217 | ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data | 2020 | VLDB |
| 7 | 1,903 | Instant Loading for Main Memory Databases | 2013 | VLDB |
| 8 | 12,051 | Vectorizing an In Situ Query Engine | 2016 | SIGMOD |
| 9 | 1,070 | NoDB: Efficient Query Execution on Raw Data Files | 2012 | SIGMOD |
| 10 | 3,062 | Adaptive Query Processing on RAW Data | 2014 | VLDB |