Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining
Summary: Proposes language-model/word-embedding–driven horizontal partitioning for frequent itemset mining: treat transactions as sentences, items as words, then cluster to form informed partitions. Goal is not just parallelism, but shrinking per-partition vocabulary/entropy to make mining scalable on large, sparse databases (e.g., graph propositionalizations). (summarized by gpt-5.4-mini on May 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Carlos Bobed (Aragon Institute of Engineering Research; University of Zaragoza)
- 2. Jorge Bernad (Aragon Institute of Engineering Research; University of Zaragoza)
- 3. Pierre Maillot (CNRS; I3S; University Cote d’Azur Inria)
BibTeX Citation
@inproceedings{bobed_sigmod24,
title = {{Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining}},
author = {Bobed, Carlos and Bernad, Jorge and Maillot, Pierre},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3654987},
url = {https://dl.acm.org/doi/10.1145/3654987},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 27 | Fast Algorithms for Mining Association Rules | 1994 | VLDB | 0.00052255472 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 558 | An Efficient Algorithm for Mining Association Rules in Large Databases | 1995 | VLDB |
| 2 | 4,640 | Scalable Parallel Data Mining for Association Rules | 1997 | SIGMOD |
| 3 | 4,547 | Mind the Gap: Large-Scale Frequent Sequence Mining | 2013 | SIGMOD |
| 4 | 5,538 | Mining Frequent Patterns with Differential Privacy | 2013 | VLDB |
| 5 | 14,018 | Communication-Efficient Distributed Mining of Association Rules | 2001 | SIGMOD |
| 6 | 161 | Mining Frequent Patterns without Candidate Generation | 2000 | SIGMOD |
| 7 | 14,090 | Towards Data Mining Benchmarking: A Test Bed for Performance Study of Frequent Pattern Mining | 2000 | SIGMOD |
| 8 | 12,922 | Parallel Mining Algorithms for Generalized Association Rules with Classification Hierarchy | 1998 | SIGMOD |
| 9 | 886 | Efficiently Mining Long Patterns from Databases | 1998 | SIGMOD |
| 10 | 9,212 | Feasible Itemset Distributions in Data Mining: Theory and Application | 2003 | PODS |