An Efficient Rigorous Approach for Identifying Statistically Significant Frequent Itemsets
Summary: Apply Chen–Stein Poisson approximation to the count of itemsets with support ≥ s to locate s* where observed counts exceed random-data expectations. Presents an efficient parametric multi-hypothesis test controlling FDR using whole-dataset counts rather than per-itemset tests; empirically validated. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Adam Kirsch (Harvard University)
- 2. Michael Mitzenmacher (Harvard University)
- 3. Andrea Pietracaprina (University of Padova)
- 4. Geppino Pucci (University of Padova)
- 5. Eli Upfal (Brown University)
- 6. Fabio Vandin (University of Padova)
BibTeX Citation
@inproceedings{kirsch_pods09,
address = {New York, NY, USA},
series = {{PODS} '09},
title = {{An Efficient Rigorous Approach for Identifying Statistically Significant Frequent Itemsets}},
url = {https://dl.acm.org/doi/10.1145/1559795.1559814},
doi = {10.1145/1559795.1559814},
booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
publisher = {Association for Computing Machinery},
author = {Kirsch, Adam and Mitzenmacher, Michael and Pietracaprina, Andrea and Pucci, Geppino and Upfal, Eli and Vandin, Fabio},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,273 | SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging | 2021 | SIGMOD | 8.8230899e-05 |
| 3,149 | Assessing and Ranking Structural Correlations in Graphs | 2011 | SIGMOD | 7.7066337e-05 |
| 12,328 | Controlling False Positives in Association Rule Mining | 2012 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 13 | Mining Association Rules between Sets of Items in Large Databases | 1993 | SIGMOD | 0.0006567919 |
| 606 | Mining Quantitative Association Rules in Large Relational Tables | 1996 | SIGMOD | 0.00015804851 |
| 3,324 | Mining Compressed Frequent-Pattern Sets | 2005 | VLDB | 7.5198821e-05 |
| 5,633 | A New Framework For Itemset Generation | 1998 | PODS | 6.1414792e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,176 | Resource-oriented Approximation for Frequent Itemset Mining from Bursty Data Streams | 2014 | SIGMOD |
| 2 | 4,092 | Traversing Itemset Lattices with Statistical Metric Pruning | 2000 | PODS |
| 3 | 885 | Finding Frequent Items in Data Streams | 2008 | VLDB |
| 4 | 2,366 | On Differentially Private Frequent Itemset Mining | 2013 | VLDB |
| 5 | 7,657 | Mining Frequent Itemsets over Uncertain Databases | 2012 | VLDB |
| 6 | 4,683 | False Positive or False Negative: Mining Frequent Itemsets from High Speed Transactional Data Streams | 2004 | VLDB |
| 7 | 12,150 | Beyond Itemsets: Mining Frequent Featuresets over Structured Items | 2015 | VLDB |
| 8 | 12,882 | Mining Frequent Itemsets Using Support Constraints | 2000 | VLDB |
| 9 | 9,212 | Feasible Itemset Distributions in Data Mining: Theory and Application | 2003 | PODS |
| 10 | 11,249 | Efficient Discovery of Significant Patterns with Few-Shot Resampling | 2024 | VLDB |