Density Biased Sampling: An Improved Method for Data Mining and Clustering
Summary: Density biased sampling under-samples dense regions and over-samples sparse ones, preserving original densities with weighted samples. Single-pass, memory-efficient algorithm; uniform sampling is a special case, with up to 6× gains on Zipf-like clusters for mining and clustering. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Christopher R. Palmer (Carnegie Mellon University)
- 2. Christos Faloutsos (Carnegie Mellon University)
BibTeX Citation
@inproceedings{palmer_sigmod00,
title = {{Density Biased Sampling: An Improved Method for Data Mining and Clustering}},
author = {Palmer, Christopher R. and Faloutsos, Christos},
series = {{SIGMOD} '00},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/342009.335384},
url = {https://dl.acm.org/doi/10.1145/342009.335384},
year = {2000}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,218 | Maintaining Variance and k–Medians over Data Stream Windows | 2003 | PODS | 8.9331834e-05 |
| 2,576 | Optimal Sampling from Sliding Windows | 2009 | PODS | 8.3964437e-05 |
| 3,881 | Using Trees to Depict a Forest | 2009 | VLDB | 7.0490549e-05 |
| 7,023 | C2P: Clustering based on Closest Pairs | 2001 | VLDB | 5.7242118e-05 |
| 7,866 | Dscaler: Synthetically Scaling A Given Relational Database | 2016 | VLDB | 5.5281143e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9 | Online Aggregation | 1997 | SIGMOD | 0.00077458002 |
| 31 | BIRCH: An Efficient Data Clustering Method for Very Large Databases | 1996 | SIGMOD | 0.00050347119 |
| 75 | Sampling-Based Estimation of the Number of Distinct Values of an Attribute | 1995 | VLDB | 0.00037277061 |
| 149 | New Sampling-Based Summary Statistics for Improving Approximate Query Answers | 1998 | SIGMOD | 0.00029226907 |
| 351 | CURE: An Efficient Clustering Algorithm for Large Databases | 1998 | SIGMOD | 0.00020424271 |
| 1,059 | On B-tree Indices for Skewed Distributions | 1992 | VLDB | 0.00012377809 |
| 1,206 | Random Sampling from Hash Files | 1990 | SIGMOD | 0.00011663837 |
| 6,053 | Modeling skewed distributions using multifractals and the '80-20 law' | 1996 | VLDB | 5.9939282e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,907 | Incremental Clustering for Mining in a Data Warehousing Environment | 1998 | VLDB |
| 2 | 10,567 | FB*: A Compact Index for Efficient and Exact Density-based Clustering | 2026 | VLDB |
| 3 | 6,974 | Uncertain Centroid based Partitional Clustering of Uncertain Data | 2012 | VLDB |
| 4 | 11,664 | Fast Density-Peaks Clustering: Multicore-based Parallelization Approach | 2021 | SIGMOD |
| 5 | 4,229 | On Biased Reservoir Sampling in the Presence of Stream Evolution | 2006 | VLDB |
| 6 | 7,903 | Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation | 2011 | SIGMOD |
| 7 | 3,020 | Dynamic Density Based Clustering | 2017 | SIGMOD |
| 8 | 2,253 | Approximation Algorithms for Clustering Uncertain Data | 2008 | PODS |
| 9 | 10,195 | Approximate DBSCAN via Density-Biased Sampling and Kernel Density Estimation | 2026 | SIGMOD |
| 10 | 2,638 | Quality and Efficiency in Kernel Density Estimates for Large Data | 2013 | SIGMOD |