Scalable Topical Phrase Mining from Text Corpora
Summary: Introduces scalable topical phrase mining by combining a phrase-mining stage with a partition-based topic model. It outperforms unigram-only methods and costly n-gram models, delivering high-quality topical phrases with negligible cost across corpora. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ahmed El-Kishky (University of Illinois Urbana-Champaign)
- 2. Yanglei Song (University of Illinois Urbana-Champaign)
- 3. Chi Wang (Microsoft)
- 4. Clare R. Voss (Army Research Laboratory)
- 5. Jiawei Han (University of Illinois Urbana-Champaign)
BibTeX Citation
@article{elkishky_vldb15,
title = {{Scalable Topical Phrase Mining from Text Corpora}},
author = {El-Kishky, Ahmed and Song, Yanglei and Wang, Chi and Voss, Clare R. and Han, Jiawei},
journal = {PVLDB},
series = {{VLDB} '15},
volume = {8},
number = {3},
pages = {305},
doi = {10.14778/2735508.2735519},
url = {https://doi.org/10.14778/2735508.2735519},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,239 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD | 5.3708698e-05 |
| 9,459 | TextCube: Automated Construction and Multidimensional Exploration | 2019 | VLDB | 5.1734318e-05 |
| 10,815 | CRAFT: Corpus Relatedness Analysis Using Fourier Transforms | 2026 | VLDB | 4.9793485e-05 |
| 12,278 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD | 4.9793485e-05 |
| 12,343 | Automatic Entity Recognition and Typing in Massive Text Data | 2016 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 29 | Fast Algorithms for Mining Association Rules | 1994 | VLDB | 0.0005121339 |
| 164 | Mining Frequent Patterns without Candidate Generation | 2000 | SIGMOD | 0.00027412227 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,340 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 2 | 5,820 | Towards the Web of Concepts: Extracting Concepts from Large Datasets | 2010 | VLDB |
| 3 | 3,939 | On the Spatiotemporal Burstiness of Terms | 2012 | VLDB |
| 4 | 4,365 | Measure-driven Keyword-Query Expansion | 2009 | VLDB |
| 5 | 2,884 | Seeking Stable Clusters in the Blogosphere | 2007 | VLDB |
| 6 | 4,736 | Scalable Ad-hoc Entity Extraction from Text Collections | 2008 | VLDB |
| 7 | 13,842 | Scalable Training of Hierarchical Topic Models | 2018 | VLDB |
| 8 | 12,330 | Topic Exploration in Spatio-Temporal Document Collections | 2016 | SIGMOD |
| 9 | 8,239 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD |
| 10 | 7,475 | Interesting-Phrase Mining for Ad-Hoc Text Analytics | 2010 | VLDB |