On Saving Outliers for Better Clustering over Noisy Data
Summary: Outlier-saving: minimally adjust erroneous values to render outliers normal, enabling clustering on the cleaned data. NP-hardness proven; bounds, a guaranteed-approximation algorithm; experiments show improved clustering and downstream tasks. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Shaoxu Song (Tsinghua University)
- 2. Fei Gao (Tsinghua University)
- 3. Ruihong Huang (Tsinghua University)
- 4. Yihan Wang (Tsinghua University)
BibTeX Citation
@inproceedings{song_sigmod21,
title = {{On Saving Outliers for Better Clustering over Noisy Data}},
author = {Song, Shaoxu and Gao, Fei and Huang, Ruihong and Wang, Yihan},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457271},
url = {https://dl.acm.org/doi/10.1145/3448016.3457271},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,659 | ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation | 2023 | VLDB | 5.2930951e-05 |
| 11,587 | Win-Win: On Simultaneous Clustering and Imputing over Incomplete Data | 2024 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 60 | The Merge/Purge Problem for Large Databases | 1995 | SIGMOD | 0.000394583 |
| 104 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00033690989 |
| 300 | OPTICS: Ordering Points To Identify the Clustering Structure | 1999 | SIGMOD | 0.00021810545 |
| 350 | Discovering Denial Constraints | 2013 | VLDB | 0.00020253521 |
| 500 | Dependencies Revisited for Improving Data Quality | 2008 | PODS | 0.00017280722 |
| 547 | ERACER: A Database Approach for Statistical Inference and Data Cleaning | 2010 | SIGMOD | 0.00016578131 |
| 652 | Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes | 2013 | SIGMOD | 0.00015121325 |
| 695 | Algorithms for Mining Distance-Based Outliers in Large Datasets | 1998 | VLDB | 0.00014702685 |
| 700 | Explaining differences in multidimensional aggregates | 1999 | VLDB | 0.00014681669 |
| 938 | Truth Finding on the Deep Web: Is the Problem Solved? | 2013 | VLDB | 0.00012973266 |
| 2,701 | Finding Intensional Knowledge of Distance-Based Outliers | 1999 | VLDB | 8.1185059e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,008 | How to Design Robust Algorithms using Noisy Comparison Oracle | 2021 | VLDB |
| 2 | 2,359 | Online Outlier Detection in Sensor Data Using Non-Parametric Models | 2006 | VLDB |
| 3 | 6,333 | Outlier-robust Clustering using Independent Components | 2008 | SIGMOD |
| 4 | 10,139 | Distance-Based Outlier Detection: Consolidation and Renewed Bearing | 2010 | VLDB |
| 5 | 695 | Algorithms for Mining Distance-Based Outliers in Large Datasets | 1998 | VLDB |
| 6 | 8,240 | Towards Metric DBSCAN: Exact, Approximate, and Streaming Algorithms | 2024 | SIGMOD |
| 7 | 4,460 | Solving k-center Clustering (with Outliers) in MapReduce and Streaming, almost as Accurately as Sequentially | 2019 | VLDB |
| 8 | 583 | Efficient Algorithms for Mining Outliers from Large Data Sets | 2000 | SIGMOD |
| 9 | 10,393 | Clustering with Set Outliers and Applications in Relational Clustering | 2026 | PODS |
| 10 | 9,751 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |