A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses
Summary: Hadoop as distributed ETL loader to Teradata EDW, leveraging HDFS for scalable, parallel loading. Polynomial-time optimal and approximate HDFS-block to Teradata-unit assignment minimizes network traffic; MapReduce enables transformation of un/ semi-structured data; experiments show gains. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yu Xu (Teradata)
- 2. Pekka Kostamaa (Teradata)
- 3. Yan Qi (Teradata)
- 4. Jian Wen (University of California Riverside)
- 5. Kevin Keliang Zhao (University of California San Diego)
BibTeX Citation
@inproceedings{xu_sigmod11,
title = {{A Hadoop Based Distributed Loading Approach to Parallel Data Warehouses}},
author = {Xu, Yu and Kostamaa, Pekka and Qi, Yan and Wen, Jian and Zhao, Kevin Keliang},
series = {{SIGMOD} '11},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/1989323.1989440},
url = {https://dl.acm.org/doi/10.1145/1989323.1989440},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,739 | Split Query Processing in Polybase | 2013 | SIGMOD | 9.8814935e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 642 | Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience | 2009 | VLDB | 0.00015395331 |
| 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.00015198804 |
| 1,791 | Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce | 2010 | VLDB | 9.7470504e-05 |
| 3,749 | Integrating Hadoop and Parallel DBMS | 2010 | SIGMOD | 7.1539955e-05 |
| 6,147 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD | 5.9601328e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 995 | Multi-Dimensional Database Allocation for Parallel Data Warehouses | 2000 | VLDB |
| 2 | 10,067 | Shared Load(ing): Efficient Bulk Loading into Optimized Storage | 2020 | CIDR |
| 3 | 8,014 | Indexing HDFS Data in PDW: Splitting the data from the index | 2014 | VLDB |
| 4 | 5,083 | Automating Distributed Tiered Storage Management in Cluster Computing | 2020 | VLDB |
| 5 | 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB |
| 6 | 12,298 | Optimization Strategies for A/B Testing on HADOOP | 2013 | VLDB |
| 7 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 8 | 11,885 | Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology | 2019 | VLDB |
| 9 | 2,159 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 10 | 3,749 | Integrating Hadoop and Parallel DBMS | 2010 | SIGMOD |