Highly-Efficient Large-Scale k-means with Individual Fairness
Summary: Introduces tilted-SSE k-means, using exponential tilting to improve individual fairness by penalizing distant assignments and reducing within-group variance. TKM/FastTKM retain Lloyd-like complexity, with stochastic estimation enabling thousand-fold speedups and lower memory at scale. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Shengkun Zhu (Wuhan University)
- 2. Jinshan Zeng (Xi'an Jiaotong University)
- 3. Yuan Sun (La Trobe University)
- 4. Sheng Wang (Wuhan University)
- 5. Yiming Wang (Wuhan University)
- 6. Yushuai Ji (Wuhan University)
- 7. Feiping Nie (Northwestern Polytechnical University)
- 8. Xiaodong Li (RMIT University)
- 9. Zhiyong Peng (Wuhan University)
BibTeX Citation
@article{zhu_vldb26,
title = {{Highly-Efficient Large-Scale k-means with Individual Fairness}},
author = {Zhu, Shengkun and Zeng, Jinshan and Sun, Yuan and Wang, Sheng and Wang, Yiming and Ji, Yushuai and Nie, Feiping and Li, Xiaodong and Peng, Zhiyong},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {5},
pages = {808--821},
doi = {10.14778/3796195.3796197},
url = {https://doi.org/10.14778/3796195.3796197},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,120 | Big Metadata: When Metadata is Big Data | 2021 | VLDB | 6.8896959e-05 |
| 5,608 | Big Graphs: Challenges and Opportunities | 2022 | VLDB | 6.1514145e-05 |
| 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB | 5.6340845e-05 |
| 7,507 | Models and Mechanisms for Spatial Data Fairness | 2023 | VLDB | 5.6029996e-05 |
| 11,420 | F3 KM: Federated, Fair, and Fast k-means | 2023 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,175 | Faster Algorithms for Fair Max-Min Diversification in Rd | 2024 | SIGMOD |
| 2 | 7,507 | Models and Mechanisms for Spatial Data Fairness | 2023 | VLDB |
| 3 | 9,575 | Local Search Methods for k-Means with Outliers | 2017 | VLDB |
| 4 | 12,764 | k-Means Projective Clustering | 2004 | PODS |
| 5 | 2,085 | Scalable K-Means++ | 2012 | VLDB |
| 6 | 5,314 | Optimal Differentially Private Algorithms for k-Means Clustering | 2018 | PODS |
| 7 | 10,959 | Federated and Balanced Clustering for High-dimensional Data | 2025 | VLDB |
| 8 | 7,360 | On the Efficiency of K-Means Clustering: Evaluation, Optimization, and Algorithm Selection | 2021 | VLDB |
| 9 | 11,161 | Efficient Algorithm for K-Multiple-Means | 2024 | SIGMOD |
| 10 | 11,420 | F3 KM: Federated, Fair, and Fast k-means | 2023 | SIGMOD |