Mining Quality Phrases from Massive Text Corpora
Summary: Proposes a scalable framework for mining quality phrases from massive text corpora by integrating phrasal segmentation with limited supervision. Demonstrates near-human phrase quality and linear time/space scalability, validated on large corpora. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jialu Liu (University of Illinois Urbana-Champaign)
- 2. Jingbo Shang (University of Illinois Urbana-Champaign)
- 3. Chi Wang (Microsoft)
- 4. Xiang Ren (University of Illinois Urbana-Champaign)
- 5. Jiawei Han (University of Illinois Urbana-Champaign)
BibTeX Citation
@inproceedings{liu_sigmod15,
title = {{Mining Quality Phrases from Massive Text Corpora}},
author = {Liu, Jialu and Shang, Jingbo and Wang, Chi and Ren, Xiang and Han, Jiawei},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2751523},
url = {https://dl.acm.org/doi/10.1145/2723372.2751523},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,459 | TextCube: Automated Construction and Multidimensional Exploration | 2019 | VLDB | 5.1734318e-05 |
| 11,538 | A Universal Sketch for Estimating Heavy Hitters and Per-Element Frequency Moments in Data Streams with Bounded Deletions | 2024 | SIGMOD | 4.9793485e-05 |
| 12,088 | GIANT: Scalable Creation of a Web-scale Ontology | 2020 | SIGMOD | 4.9793485e-05 |
| 12,278 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD | 4.9793485e-05 |
| 12,343 | Automatic Entity Recognition and Typing in Massive Text Data | 2016 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,536 | Multidimensional Content eXploration | 2008 | VLDB | 7.2241245e-05 |
| 5,820 | Towards the Web of Concepts: Extracting Concepts from Large Datasets | 2010 | VLDB | 5.9832494e-05 |
| 7,475 | Interesting-Phrase Mining for Ad-Hoc Text Analytics | 2010 | VLDB | 5.5164354e-05 |
| 9,498 | Scalable Topical Phrase Mining from Text Corpora | 2015 | VLDB | 5.1708619e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,529 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 2 | 12,259 | Scalable Semantic Querying of Text | 2018 | VLDB |
| 3 | 2,880 | Expressive and Flexible Access to Web-Extracted Data: A Keyword-based Structured Query Language | 2010 | SIGMOD |
| 4 | 12,340 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 5 | 7,417 | Cardinality Estimation of Approximate Substring Queries using Deep Learning | 2022 | VLDB |
| 6 | 12,460 | Mining Latent Entity Structures from Massive Unstructured and Interconnected Data | 2014 | SIGMOD |
| 7 | 6,900 | Keyword Query Cleaning | 2008 | VLDB |
| 8 | 12,278 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 9 | 9,498 | Scalable Topical Phrase Mining from Text Corpora | 2015 | VLDB |
| 10 | 7,475 | Interesting-Phrase Mining for Ad-Hoc Text Analytics | 2010 | VLDB |