DBScholar

Back to papers

Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology

Summary: Hybrid SQL analytics unites MapReduce/Hadoop with parallel DBMS (Greenplum/Vertica); HadoopDB prototype shows strong SQL performance, scalability, and fault tolerance. Tracing a decade from research to Hadapt/Teradata, the paper surveys open-source ecosystems sustaining the integrated data-processing and DBMS paradigm. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h529c3d20aee33311
Venue
VLDB
Year
2019
Pagerank
4.9769913e-05
Overall Rank
12,191 | 18.07%
DOI
10.14778/3352063.3352145
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@article{abouzied_vldb19,
        title = {{Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology}},
        author = {Abouzied, Azza and Abadi, Daniel J. and Bajda-Pawlikowski, Kamil and Silberschatz, Avi},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {12},
        pages = {2290--2299},
        doi = {10.14778/3352063.3352145},
        url = {https://doi.org/10.14778/3352063.3352145},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 20 of 20 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010515896
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050475202
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049821554
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
48 Weaving Relations for Cache Performance 2001 VLDB 0.00043795812
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004314366
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031099083
327 Impala: A Modern, Open-Source SQL Engine for Hadoop 2015 CIDR 0.00020942751
379 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00019507406
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018331051
512 Scalable SPARQL Querying of Large RDF Graphs 2011 VLDB 0.00017053842
1,745 Sinew: A SQL System for Multi-Structured Data 2014 SIGMOD 9.7324334e-05
1,881 Ground: A Data Context Service 2017 CIDR 9.4372419e-05
2,176 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.9150466e-05
2,379 LEOPARD: Lightweight Edge-Oriented Partitioning and Replication for Dynamic Graphs 2016 VLDB 8.5470334e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
3,446 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing 2019 SIGMOD 7.2908619e-05
3,891 Apache Tez: A Unifying Framework for Modeling and Building Data Processing Applications 2015 SIGMOD 6.9400948e-05
4,091 Automatic Generation of Normalized Relational Schemas from Nested Key-Value Data 2016 SIGMOD 6.8116205e-05
Previous Page 1 / 1 Next

Semantically Similar Papers