Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)
Summary: RAMBO uses a Repeated And Merged Bloom Filter to turn genome search into count-min style set membership tests, delivering zero false negatives and a small index. Streaming updates and indexing 170 TB in 9 hours, beating COBS, HowDeSBT, and SSBT. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Gaurav Gupta (Rice University)
- 2. Minghao Yan (Rice University)
- 3. Benjamin Coleman (Rice University)
- 4. Bryce Kille (Rice University)
- 5. R. A. Leo Elworth (Rice University)
- 6. Tharun Medini (Rice University)
- 7. Todd Treangen (Rice University)
- 8. Anshumali Shrivastava (Rice University)
BibTeX Citation
@inproceedings{gupta_sigmod21,
title = {{Fast Processing and Querying of 170TB of Genomics Data via a Repeated And Merged BloOm Filter (RAMBO)}},
author = {Gupta, Gaurav and Yan, Minghao and Coleman, Benjamin and Kille, Bryce and Elworth, R. A. Leo and Medini, Tharun and Treangen, Todd and Shrivastava, Anshumali},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457333},
url = {https://dl.acm.org/doi/10.1145/3448016.3457333},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,474 | GTS: GPU-based Tree Index for Fast Similarity Search | 2024 | SIGMOD | 5.6093517e-05 |
| 7,716 | Double-Anonymous Sketch: Achieving Top-K-fairness for Finding Global Top-K Frequent Items | 2023 | SIGMOD | 5.5605526e-05 |
| 11,488 | ChainDash: An Ad-Hoc Blockchain Data Analytics System | 2023 | VLDB | 5.093636e-05 |
| 11,572 | New Wine in an Old Bottle: Data-Aware Hash Functions for Bloom Filters | 2022 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 862 | Spectral Bloom Filters | 2003 | SIGMOD | 0.00013532857 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,026 | A New Approach for Processing Ranked Subsequence Matching Based on Ranked Union | 2011 | SIGMOD |
| 2 | 8,429 | Building Highly-Optimized, Low-Latency Pipelines for Genomic Data Analysis | 2015 | CIDR |
| 3 | 8,612 | Building Fast and Compact Sketches for Approximately Multi-Set Multi-Membership Querying | 2021 | SIGMOD |
| 4 | 10,086 | Efficient and Effective KNN Sequence Search with Approximate n-grams | 2014 | VLDB |
| 5 | 2,008 | A Database Index to Large Biological Sequences | 2001 | VLDB |
| 6 | 13,656 | Memory Efficient Minimum Substring Partitioning | 2013 | VLDB |
| 7 | 6,496 | Reference-Based Indexing of Sequence Databases | 2006 | VLDB |
| 8 | 5,242 | Serial and Parallel Methods for I/O Efficient Suffix Tree Construction | 2009 | SIGMOD |
| 9 | 12,283 | RCSI: Scalable similarity search in thousand(s) of genomes | 2013 | VLDB |
| 10 | 3,077 | Genome-scale Disk-based Suffix Tree Indexing | 2007 | SIGMOD |