Data Stream Clustering: An In-depth Empirical Study
Summary: Empirical DSC study across four design axes: data summarization, windowing, outlier detection, and offline refinement; implemented from scratch and tested on real and synthetic streams. Introduces Benne, a tunable hybrid that can boost accuracy or efficiency by mixing design choices. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xin Wang (Ohio State University; Sichuan University)
- 2. Zhengru Wang (Huazhong University of Science and Technology; NVIDIA)
- 3. Zhenyu Wu (University of Manchester)
- 4. Shuhao Zhang (Singapore Institute of Technology)
- 5. Xuanhua Shi (Huazhong University of Science and Technology)
- 6. Li Lu (Sichuan University)
BibTeX Citation
@inproceedings{wang_sigmod23,
title = {{Data Stream Clustering: An In-depth Empirical Study}},
author = {Wang, Xin and Wang, Zhengru and Wu, Zhenyu and Zhang, Shuhao and Shi, Xuanhua and Lu, Li},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3589307},
url = {https://dl.acm.org/doi/10.1145/3589307},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,961 | BURST: Rendering Clustering Techniques Suitable for Evolving Streams | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 31 | BIRCH: An Efficient Data Clustering Method for Very Large Databases | 1996 | SIGMOD | 0.00050347119 |
| 907 | A Framework for Clustering Evolving Data Streams | 2003 | VLDB | 0.00013309819 |
| 5,335 | Clustering Stream Data by Exploring the Evolution of Density Mountain | 2018 | VLDB | 6.2618226e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,978 | Time-Series Clustering: A Comprehensive Study of Data Mining, Machine Learning, and Deep Learning Methods | 2025 | VLDB |
| 2 | 4,366 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB |
| 3 | 3,958 | Continuous Outlier Detection in Data Streams: An Extensible Framework and State-Of-The-Art Algorithms | 2013 | SIGMOD |
| 4 | 1,274 | Querying and Mining Data Streams: You Only Get One Look | 2002 | SIGMOD |
| 5 | 10,497 | Scalable Clustering Over High Dimensional Vector Streams | 2026 | SIGMOD |
| 6 | 8,881 | Summarization and Matching of Density-Based Clusters in Streaming Environments | 2012 | VLDB |
| 7 | 8,070 | Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms | 2024 | SIGMOD |
| 8 | 9,272 | A Framework for Projected Clustering of High Dimensional Data Streams | 2004 | VLDB |
| 9 | 907 | A Framework for Clustering Evolving Data Streams | 2003 | VLDB |
| 10 | 5,335 | Clustering Stream Data by Exploring the Evolution of Density Mountain | 2018 | VLDB |