Interesting-Phrase Mining for Ad-Hoc Text Analytics
Summary: Introduces a phrase-centric framework for ad-hoc text analytics, prioritizing multi-word phrases that are frequent in a subset yet rare in the full corpus. Develops preprocessing, indexing, and top-k search methods for scalable discovery, validated on a large NYT corpus. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Srikanta Bedathur
- 2. Klaus Berberich
- 3. Jens Dittrich
- 4. Nikos Mamoulis
- 5. Gerhard Weikum
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,913 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD | 4.6139203e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 182 | Mining Frequent Patterns without Candidate Generation | 2000 | SIGMOD | 0.00036955562 |
| 2,171 | BlogScope: A System for Online Analysis of High Volume Text Streams | 2007 | VLDB | 9.3805639e-05 |
| 3,261 | Multidimensional Content eXploration | 2008 | VLDB | 7.3088095e-05 |
| 4,690 | Multi-Structural Databases | 2005 | PODS | 5.9898251e-05 |
| 6,368 | Efficient Implementation of Large-Scale Multi-Structural Databases | 2005 | VLDB | 5.0886655e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,979 | Mining Latent Entity Structures from Massive Unstructured and Interconnected Data | 2014 | SIGMOD | 4.1905499e-05 |
| 11,842 | Topic Exploration in Spatio-Temporal Document Collections | 2016 | SIGMOD | 4.1905499e-05 |
| 9,310 | Phrase Matching in XML | 2003 | VLDB | 4.3536718e-05 |
| 13,414 | NewsNetExplorer: Automatic Construction and Exploration of News Information Networks | 2014 | SIGMOD | - |
| 11,852 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD | 4.1905499e-05 |
| 7,890 | Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation | 2011 | SIGMOD | 4.6205184e-05 |
| 5,382 | Scalable Ad-hoc Entity Extraction from Text Collections | 2008 | VLDB | 5.5358382e-05 |
| 11,783 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD | 4.1905499e-05 |
| 7,913 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD | 4.6139203e-05 |
| 11,962 | Scalable Topical Phrase Mining from Text Corpora | 2015 | VLDB | 4.1905499e-05 |