DBScholar

Back to papers

Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology

Summary: Hybrid SQL analytics unites MapReduce/Hadoop with parallel DBMS (Greenplum/Vertica); HadoopDB prototype shows strong SQL performance, scalability, and fault tolerance. Tracing a decade from research to Hadapt/Teradata, the paper surveys open-source ecosystems sustaining the integrated data-processing and DBMS paradigm. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h529c3d20aee33311
Venue
VLDB
Year
2019
Pagerank
4.9793485e-05
Overall Rank
12,185 | 18.08%
DOI
10.14778/3352063.3352145

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{abouzied_vldb19,
        title = {{Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology}},
        author = {Abouzied, Azza and Abadi, Daniel J. and Bajda-Pawlikowski, Kamil and Silberschatz, Avi},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {12},
        pages = {2290--2299},
        doi = {10.14778/3352063.3352145},
        url = {https://doi.org/10.14778/3352063.3352145},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
48 Weaving Relations for Cache Performance 2001 VLDB 0.00043805923
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.0002095191
379 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00019514689
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
511 Scalable SPARQL Querying of Large RDF Graphs 2011 VLDB 0.00017061883
1,743 Sinew: A SQL System for Multi-Structured Data 2014 SIGMOD 9.7368807e-05
1,880 Ground: A Data Context Service 2017 CIDR 9.4415148e-05
2,173 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.91924e-05
2,377 LEOPARD: Lightweight Edge-Oriented Partitioning and Replication for Dynamic Graphs 2016 VLDB 8.5510365e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
3,446 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing 2019 SIGMOD 7.2942885e-05
3,891 Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications 2015 SIGMOD 6.9432955e-05
4,088 Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data 2016 SIGMOD 6.8148416e-05
Previous Page 1 / 1 Next

Semantically Similar Papers