Interesting-Phrase Mining for Ad-Hoc Text Analytics
Summary: Introduces a phrase-centric framework for ad-hoc text analytics, prioritizing multi-word phrases that are frequent in a subset yet rare in the full corpus. Develops preprocessing, indexing, and top-k search methods for scalable discovery, validated on a large NYT corpus. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Srikanta Bedathur (Max Planck Institute)
- 2. Klaus Berberich (Max Planck Institute)
- 3. Jens Dittrich (Saarland University)
- 4. Nikos Mamoulis (Max Planck Institute; University of Hong Kong)
- 5. Gerhard Weikum (Max Planck Institute)
BibTeX Citation
@article{bedathur_vldb10,
title = {{Interesting-Phrase Mining for Ad-Hoc Text Analytics}},
author = {Bedathur, Srikanta and Berberich, Klaus and Dittrich, Jens and Mamoulis, Nikos and Weikum, Gerhard},
journal = {PVLDB},
series = {{VLDB} '10},
volume = {3},
number = {1},
pages = {1348--1359},
doi = {10.14778/1920841.1921007},
url = {https://doi.org/10.14778/1920841.1921007},
year = {2010}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,067 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD | 5.4941436e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 161 | Mining Frequent Patterns without Candidate Generation | 2000 | SIGMOD | 0.00027981772 |
| 2,703 | BlogScope: A System for Online Analysis of High Volume Text Streams | 2007 | VLDB | 8.2340288e-05 |
| 3,467 | Multidimensional Content eXploration | 2008 | VLDB | 7.3898968e-05 |
| 4,392 | Multi-Structural Databases | 2005 | PODS | 6.7288249e-05 |
| 6,452 | Efficient Implementation of Large-Scale Multi-Structural Databases | 2005 | VLDB | 5.8751532e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,169 | Mining Latent Entity Structures from Massive Unstructured and Interconnected Data | 2014 | SIGMOD |
| 2 | 2,833 | Seeking Stable Clusters in the Blogosphere | 2007 | VLDB |
| 3 | 9,600 | Phrase Matching in XML | 2003 | VLDB |
| 4 | 13,602 | NewsNetExplorer: Automatic Construction and Exploration of News Information Networks | 2014 | SIGMOD |
| 5 | 7,903 | Mining a Search Engine’s Corpus: Efficient Yet Unbiased Sampling and Aggregate Estimation | 2011 | SIGMOD |
| 6 | 12,045 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 7 | 4,992 | Scalable Ad-hoc Entity Extraction from Text Collections | 2008 | VLDB |
| 8 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 9 | 8,067 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD |
| 10 | 12,152 | Scalable Topical Phrase Mining from Text Corpora | 2015 | VLDB |