Clydesdale: Structured Data Processing on Hadoop
Summary: Clydesdale, a Hadoop-based prototype for structured data processing, achieves major performance gains without changing MapReduce. By fusing DB techniques with Hadoop and exposing ClyQL, a Scala DSL for star-joins, it delivers ~38x faster star-schema workloads than Hive. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Andrey Balmin (IBM)
- 2. Tim Kaldewey (IBM)
- 3. Sandeep Tata (IBM)
BibTeX Citation
@inproceedings{balmin_sigmod12,
title = {{Clydesdale: Structured Data Processing on Hadoop}},
author = {Balmin, Andrey and Kaldewey, Tim and Tata, Sandeep},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213938},
url = {https://dl.acm.org/doi/10.1145/2213836.2213938},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,230 | Can the Elephants Handle the NoSQL Onslaught? | 2012 | VLDB | 6.8178353e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12 | C-Store: A Column-oriented DBMS | 2005 | VLDB | 0.00069513174 |
| 29 | Database Architecture Optimized for the New Bottleneck: Memory Access | 1999 | VLDB | 0.00052093615 |
| 44 | A Comparison of Approaches to Large-Scale Data Analysis | 2009 | SIGMOD | 0.00046055057 |
| 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB | 0.00031680027 |
| 632 | Adaptive Aggregation on Chip Multiprocessors | 2007 | VLDB | 0.00015575286 |
| 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.00015198804 |
| 2,849 | Column-Oriented Storage Techniques for MapReduce | 2011 | VLDB | 8.053191e-05 |
| 2,942 | Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework | 2011 | SIGMOD | 7.9358593e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,276 | CARTILAGE: Adding Flexibility to the Hadoop Skeleton | 2013 | SIGMOD |
| 2 | 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB |
| 3 | 2,159 | Efficient Processing of Data Warehousing Queries in a Split Execution Environment | 2011 | SIGMOD |
| 4 | 12,170 | Palette: Enabling Scalable Analytics for Big-Memory, Multicore Machines | 2014 | SIGMOD |
| 5 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 6 | 5,902 | A Demonstration of ST-Hadoop: A MapReduce Framework for Big Spatio-temporal Data | 2017 | VLDB |
| 7 | 6,147 | HadoopDB in Action: Building Real World Applications | 2010 | SIGMOD |
| 8 | 9,510 | Efficient Big Data Processing in Hadoop MapReduce | 2012 | VLDB |
| 9 | 9,829 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive | 2022 | SIGMOD |
| 10 | 12,146 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB |