DBScholar

Back to papers

Mining Statistically Significant Substrings using the Chi-Square Statistic

Summary: Defines the most significant substring as the contiguous region whose letter distribution maximally departs from a Bernoulli model under chi-square. Gives an O(n^{3/2})-time algorithm w.h.p., plus top-t, threshold, and minimum-length variants. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10532
Venue
VLDB
Year
2012
Pagerank
5.4904274e-05
Overall Rank
8,087 | 44.52%
DOI
10.14778/2336664.2336673

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{sachan_vldb12,
        title = {{Mining Statistically Significant Substrings using the Chi-Square Statistic}},
        author = {Sachan, Mayank and Bhattacharya, Arnab},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {10},
        pages = {1052--1063},
        doi = {10.14778/2336664.2336673},
        url = {https://doi.org/10.14778/2336664.2336673},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Rank Citing Paper Year Venue Pagerank
7,231 Mining Statistically Significant Connected Subgraphs in Vertex Labeled Graphs 2014 SIGMOD 5.6664362e-05
7,740 Oracle Workload Intelligence 2015 SIGMOD 5.5550801e-05
9,665 ChiSeL: Graph Similarity Search using Chi-Squared Statistics in Large Probabilistic Graphs 2020 VLDB 5.2406724e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 0 of 0 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Semantically Similar Papers