Parallel In-Situ Data Processing with Speculative Loading
Summary: SCANRAW is a parallel in-situ operator for raw files, merging loading with external tables to preserve zero time-to-query. It employs a parallel super-scalar pipeline with speculative loading to overlap queries and conversion, balancing CPU and I/O. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yu Cheng (University of California Merced)
- 2. Florin Rusu (University of California Merced)
BibTeX Citation
@inproceedings{cheng_sigmod14,
title = {{Parallel In-Situ Data Processing with Speculative Loading}},
author = {Cheng, Yu and Rusu, Florin},
series = {{SIGMOD} '14},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2588555.2593673},
url = {https://dl.acm.org/doi/10.1145/2588555.2593673},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 19 of 19 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 53 | Eddies: Continuously Adaptive Query Processing | 2000 | SIGMOD | 0.00040860054 |
| 887 | Requirements for Science Data Bases and SciDB | 2009 | CIDR | 0.00013258595 |
| 1,040 | The DataPath System: A Data-Centric Analytic Processing Engine for Large Data Warehouses | 2010 | SIGMOD | 0.00012364063 |
| 1,069 | NoDB: Efficient Query Execution on Raw Data Files | 2012 | SIGMOD | 0.00012185253 |
| 1,836 | Merging What's Cracked, Cracking What's Merged: Adaptive Indexing in Main-Memory Column-Stores | 2011 | VLDB | 9.535551e-05 |
| 1,863 | The Researcher's Guide to the Data Deluge: Querying a Scientific Database in Just a Few Seconds | 2011 | VLDB | 9.4846052e-05 |
| 1,925 | Here are my Data Files. Here are my Queries. Where are my Results? | 2011 | CIDR | 9.3670488e-05 |
| 1,931 | Instant Loading for Main Memory Databases | 2013 | VLDB | 9.3474242e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,305 | Multi-dimensional Resource Scheduling for Parallel Queries | 1996 | SIGMOD |
| 2 | 1,925 | Here are my Data Files. Here are my Queries. Where are my Results? | 2011 | CIDR |
| 3 | 12,752 | Ten Thousand SQLs: Parallel Keyword Queries Computing | 2010 | VLDB |
| 4 | 6,288 | Elastic Pipelining in an In-Memory Database Cluster | 2016 | SIGMOD |
| 5 | 3,445 | Slalom: Coasting Through Raw Data via Adaptive Partitioning and Indexing | 2017 | VLDB |
| 6 | 6,063 | ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data | 2020 | VLDB |
| 7 | 1,931 | Instant Loading for Main Memory Databases | 2013 | VLDB |
| 8 | 12,346 | Vectorizing an In Situ Query Engine | 2016 | SIGMOD |
| 9 | 1,069 | NoDB: Efficient Query Execution on Raw Data Files | 2012 | SIGMOD |
| 10 | 3,069 | Adaptive Query Processing on RAW Data | 2014 | VLDB |