Rethinking Data-Intensive Science Using Scalable Analytics Systems
Summary: Maps scientific pipelines to commodity big-data platforms (Spark/Parquet) for scalable data-intensive science. ADAM delivers 28x genomics speedup and 63% cost savings; 2.8–8.9x astronomy gains, techniques for efficient analyses on big-data platforms. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Frank Austin Nothaft (University of California Berkeley)
- 2. Matt Massie (University of California Berkeley)
- 3. Timothy Danford (Genomebridge; University of California Berkeley)
- 4. Zhao Zhang (University of California Berkeley)
- 5. Uri Laserson (Cloudera)
- 6. Carl Yeksigian (Genomebridge)
- 7. Jey Kottalam (University of California Berkeley)
- 8. Arun Ahuja (Icahn School of Medicine at Mount Sinai)
- 9. Jeff Hammerbacher (Cloudera; Icahn School of Medicine at Mount Sinai)
- 10. Michael Linderman (Icahn School of Medicine at Mount Sinai)
- 11. Michael J. Franklin (University of California Berkeley)
- 12. Anthony D. Joseph (University of California Berkeley)
- 13. David A. Patterson (University of California Berkeley)
BibTeX Citation
@inproceedings{nothaft_sigmod15,
title = {{Rethinking Data-Intensive Science Using Scalable Analytics Systems}},
author = {Nothaft, Frank Austin and Massie, Matt and Danford, Timothy and Zhang, Zhao and Laserson, Uri and Yeksigian, Carl and Kottalam, Jey and Ahuja, Arun and Hammerbacher, Jeff and Linderman, Michael and Franklin, Michael J. and Joseph, Anthony D. and Patterson, David A.},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2742787},
url = {https://dl.acm.org/doi/10.1145/2723372.2742787},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 520 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores | 2020 | VLDB | 0.00017136828 |
| 3,208 | ForkBase: An Efficient Storage Engine for Blockchain and Forkable Applications | 2018 | VLDB | 7.635872e-05 |
| 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB | 7.436229e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 51 | Dremel: Interactive Analysis of Web-Scale Datasets | 2010 | VLDB | 0.0004291425 |
| 60 | Integrating Compression and Execution in Column-Oriented Database Systems | 2006 | SIGMOD | 0.0003955489 |
| 186 | The Vertica Analytic Database: C-Store 7 Years Later | 2012 | VLDB | 0.00026182534 |
| 239 | Overview of SciDB: Large Scale Array Storage, Processing and Analysis | 2010 | SIGMOD | 0.00023674329 |
| 330 | Impala: A Modern, Open-Source SQL Engine for Hadoop | 2015 | CIDR | 0.0002104801 |
| 1,067 | Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce | 2013 | VLDB | 0.00012327784 |
| 2,459 | WHAM: A High-throughput Sequence Alignment Method | 2011 | SIGMOD | 8.550768e-05 |
| 8,429 | Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis | 2015 | CIDR | 5.4263087e-05 |
| 8,430 | A Demonstration of Iterative Parallel Array Processing in Support of Telescope Image Analysis | 2013 | VLDB | 5.4263087e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,850 | Bridging the Gap Between HPC and Big Data Frameworks | 2017 | VLDB |
| 2 | 11,994 | Massively Parallel Processing of Whole Genome Sequence Data: An In-Depth Performance Study | 2017 | SIGMOD |
| 3 | 12,484 | Data Management for High-Throughput Genomics | 2009 | CIDR |
| 4 | 6,667 | Managing Data from High-Throughput Genomic Processing: A Case Study | 2004 | VLDB |
| 5 | 11,874 | I Can't Believe It's Not (Only) Software! Bionic Distributed Storage for Parquet Files | 2019 | VLDB |
| 6 | 13,557 | Big Data Science Needs Big Data Middleware | 2015 | CIDR |
| 7 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 8 | 4,224 | Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads | 2017 | VLDB |
| 9 | 2,389 | Parallel Data Analysis Directly on Scientific File Formats | 2014 | SIGMOD |
| 10 | 8,429 | Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis | 2015 | CIDR |