Massively Parallel Processing of Whole Genome Sequence Data: An In-Depth Performance Study
Summary: Gesall: a massively parallel platform for genome analysis using Wrapper Technology and GDPT to wrap programs unmodified. Genomics workloads show super-linear and sublinear speedups, highlighting data-management questions at the genomics frontier. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Abhishek Roy (University of Massachusetts Amherst)
- 2. Yanlei Diao (University of Massachusetts Amherst)
- 3. Uday Evani (New York Genome Center)
- 4. Avinash Abhyankar (New York Genome Center)
- 5. Clinton Howarth (New York Genome Center)
- 6. Rémi Le Priol (Ecole Polytechnique)
- 7. Toby Bloom (New York Genome Center)
BibTeX Citation
@inproceedings{roy_sigmod17,
title = {{Massively Parallel Processing of Whole Genome Sequence Data: An In-Depth Performance Study}},
author = {Roy, Abhishek and Diao, Yanlei and Evani, Uday and Abhyankar, Avinash and Howarth, Clinton and Le Priol, Rémi and Bloom, Toby},
series = {{SIGMOD} '17},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3035918.3064048},
url = {https://dl.acm.org/doi/10.1145/3035918.3064048},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD | 8.8398946e-05 |
| 2,459 | WHAM: A High-throughput Sequence Alignment Method | 2011 | SIGMOD | 8.550768e-05 |
| 2,949 | GenBase: A Complex Analytics Genomics Benchmark | 2014 | SIGMOD | 7.9281519e-05 |
| 8,429 | Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis | 2015 | CIDR | 5.4263087e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,297 | RACE: A Scalable and Elastic Parallel System for Discovering Repeats in Very Long Sequences | 2013 | VLDB |
| 2 | 12,801 | Genomics Algebra: A New, Integrating Data Model, Language, and Tool for Processing and Querying Genomic Information | 2003 | CIDR |
| 3 | 5,242 | Serial and Parallel Methods for I/O Efficient Suffix Tree Construction | 2009 | SIGMOD |
| 4 | 2,949 | GenBase: A Complex Analytics Genomics Benchmark | 2014 | SIGMOD |
| 5 | 14,068 | Information Management for Genome Level Bioinformatics | 2001 | VLDB |
| 6 | 3,552 | Rethinking Data-Intensive Science Using Scalable Analytics Systems | 2015 | SIGMOD |
| 7 | 12,484 | Data Management for High-Throughput Genomics | 2009 | CIDR |
| 8 | 6,667 | Managing Data from High-Throughput Genomic Processing: A Case Study | 2004 | VLDB |
| 9 | 12,335 | Massive Genomic Data Processing and Deep Analysis | 2012 | VLDB |
| 10 | 8,429 | Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis | 2015 | CIDR |