Clydesdale: Structured Data Processing on Hadoop
Summary: Clydesdale, a Hadoop-based prototype for structured data processing, achieves major performance gains without changing MapReduce. By fusing DB techniques with Hadoop and exposing ClyQL, a Scala DSL for star-joins, it delivers ~38x faster star-schema workloads than Hive. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Andrey Balmin (IBM)
- 2. Tim Kaldewey (IBM)
- 3. Sandeep Tata (IBM)
BibTeX Citation
@inproceedings{balmin_sigmod12,
title = {{Clydesdale: Structured Data Processing on Hadoop}},
author = {Balmin, Andrey and Kaldewey, Tim and Tata, Sandeep},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213938},
url = {https://dl.acm.org/doi/10.1145/2213836.2213938},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,289 | Can the Elephants Handle the NoSQL Onslaught? | 2012 | VLDB | 6.6838669e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12 | C-Store: A Column-oriented DBMS | 2005 | VLDB | 0.00068998927 |
| 27 | Database Architecture Optimized for the New Bottleneck: Memory Access | 1999 | VLDB | 0.0005158963 |
| 43 | A Comparison of Approaches to Large-Scale Data Analysis | 2009 | SIGMOD | 0.00045546775 |
| 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB | 0.000311132 |
| 626 | Adaptive Aggregation on Chip Multiprocessors | 2007 | VLDB | 0.00015473276 |
| 673 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.0001488755 |
| 2,903 | Column-Oriented Storage Techniques for MapReduce | 2011 | VLDB | 7.8810111e-05 |
| 3,007 | Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework | 2011 | SIGMOD | 7.7607173e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB |
| 2 | 8,442 | CARTILAGE: Adding Flexibility to the Hadoop Skeleton | 2013 | SIGMOD |
| 3 | 2,173 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 4 | 12,461 | Palette: Enabling Scalable Analytics for Big-Memory, Multicore Machines | 2014 | SIGMOD |
| 5 | 1,884 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 6 | 5,959 | A Demonstration of ST-Hadoop: A MapReduce Framework for Big Spatio-temporal Data | 2017 | VLDB |
| 7 | 6,276 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD |
| 8 | 9,693 | Efficient Big Data Processing in Hadoop MapReduce | 2012 | VLDB |
| 9 | 10,015 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive | 2022 | SIGMOD |
| 10 | 12,438 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB |