Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing
Summary: Introduces Parallel LSH, a cache-conscious, multicore/ multinode variant with fast construction, duplicate elimination, insert-optimized storage, and expiration for streaming data. Scales similarity search to >1B tweets at 1–2.5 ms/query, up to 8.3× faster than basic LSH. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Narayanan Sundaram (Intel)
- 2. Aizana Turmukhametova (Massachusetts Institute of Technology)
- 3. Nadathur Satish (Intel)
- 4. Todd Mostak (Massachusetts Institute of Technology)
- 5. Piotr Indyk (Massachusetts Institute of Technology)
- 6. Samuel Madden (Massachusetts Institute of Technology)
- 7. Pradeep Dubey (Intel)
BibTeX Citation
@article{sundaram_vldb13,
title = {{Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing}},
author = {Sundaram, Narayanan and Turmukhametova, Aizana and Satish, Nadathur and Mostak, Todd and Indyk, Piotr and Madden, Samuel and Dubey, Pradeep},
journal = {PVLDB},
series = {{VLDB} '13},
volume = {6},
number = {14},
pages = {1930--1941},
doi = {10.14778/2556549.2556574},
url = {https://doi.org/10.14778/2556549.2556574},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 209 | Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs | 2009 | VLDB | 0.00024932174 |
| 219 | A Study of Index Structures for Main Memory Database Management Systems | 1986 | VLDB | 0.00024293529 |
| 221 | Robust and Fast Similarity Search for Moving Object Trajectories | 2005 | SIGMOD | 0.00024224879 |
| 591 | Substructure Similarity Search in Graph Databases | 2005 | SIGMOD | 0.0001603683 |
| 678 | Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort | 2010 | SIGMOD | 0.00015061068 |
| 748 | Effective Keyword Search in Relational Databases | 2006 | SIGMOD | 0.00014381224 |
| 2,459 | WHAM: A High-throughput Sequence Alignment Method | 2011 | SIGMOD | 8.550768e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,358 | Parallel Index-Based Structural Graph Clustering and Its Approximation | 2021 | SIGMOD |
| 2 | 581 | Quality and Efficiency in High Dimensional Nearest Neighbor Search | 2009 | SIGMOD |
| 3 | 3,866 | Fast Parallel Similarity Search in Multimedia Databases | 1997 | SIGMOD |
| 4 | 2,519 | Similarity search in the blink of an eye with compressed indices | 2023 | VLDB |
| 5 | 10,265 | LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH | 2026 | SIGMOD |
| 6 | 8,613 | Bidirectionally Densifying LSH Sketches with Empty Bins | 2021 | SIGMOD |
| 7 | 6,597 | Similarity Search and Locality Sensitive Hashing using Ternary Content Addressable Memories | 2010 | SIGMOD |
| 8 | 5,666 | Smooth Tradeoffs between Insert and Query Complexity in Nearest Neighbor Search | 2015 | PODS |
| 9 | 287 | Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search | 2007 | VLDB |
| 10 | 21 | Similarity Search in High Dimensions via Hashing | 1999 | VLDB |