Approximate Matching of Hierarchical Data Using pq-Grams
Summary: Approximate matching of hierarchical data via pq-grams for autonomous sources. The pq-gram distance provides an efficient, scalable approximation of tree edit distance for ordered labeled trees, enabling near-matches in hierarchical records (e.g., addresses) and is validated with synthetic and real data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Nikolaus Augsten (Free University of Bolzano)
- 2. Michael Böhlen (Free University of Bolzano)
- 3. Johann Gamper (Free University of Bolzano)
BibTeX Citation
@article{augsten_vldb05,
title = {{Approximate Matching of Hierarchical Data Using pq-Grams}},
author = {Augsten, Nikolaus and Böhlen, Michael and Gamper, Johann},
journal = {PVLDB},
series = {{VLDB} '05},
year = {2005}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,133 | Comparing Stars: On Approximating Graph Edit Distance | 2009 | VLDB | 0.00011892544 |
| 6,974 | A Scalable Index for Top-k Subtree Similarity Queries | 2019 | SIGMOD | 5.6277375e-05 |
| 7,625 | SyncSignature: A Simple, Efficient, Parallelizable Framework for Tree Similarity Joins | 2023 | VLDB | 5.4806154e-05 |
| 7,705 | An Incrementally Maintainable Index for Approximate Lookups in Hierarchical Data | 2006 | VLDB | 5.472514e-05 |
| 11,347 | Extensible and Robust Evaluation of Similarity Queries | 2025 | VLDB | 4.9769913e-05 |
| 12,583 | Synthetising Changes in XML Documents as PULs | 2013 | VLDB | 4.9769913e-05 |
| 12,846 | The Power of Two Min-Hashes for Similarity Search among Hierarchical Data Objects | 2008 | PODS | 4.9769913e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB | 0.00033040246 |
| 166 | On Supporting Containment Queries in Relational Database Management Systems | 2001 | SIGMOD | 0.00027241844 |
| 179 | Holistic Twig Joins: Optimal XML Pattern Matching | 2002 | SIGMOD | 0.00026617591 |
| 1,235 | A Comprehensive XQuery to SQL Translation using Dynamic Interval Encoding | 2003 | SIGMOD | 0.00011400492 |
| 1,532 | Change Detection in Hierarchically Structured Information | 1996 | SIGMOD | 0.00010333424 |
| 2,746 | Holistic Twig Joins on Indexed XML Documents | 2003 | VLDB | 8.0587951e-05 |
| 2,747 | Approximate XML Joins | 2002 | SIGMOD | 8.0576613e-05 |
| 3,248 | Approximate XML Query Answers | 2004 | SIGMOD | 7.4912478e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,701 | TreeSpan: Efficiently Computing Similarity All-Matching | 2012 | SIGMOD |
| 2 | 2,969 | Extending Q-Grams to Estimate Selectivity of String Matching with Low Edit Distance | 2007 | VLDB |
| 3 | 12,846 | The Power of Two Min-Hashes for Similarity Search among Hierarchical Data Objects | 2008 | PODS |
| 4 | 108 | Approximate String Joins in a Database (Almost) for Free | 2001 | VLDB |
| 5 | 12,065 | Approximate Pattern Matching in Massive Graphs with Precision and Recall Guarantees | 2020 | SIGMOD |
| 6 | 7,846 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 7 | 2,747 | Approximate XML Joins | 2002 | SIGMOD |
| 8 | 6,140 | Scaling Similarity Joins over Tree-Structured Data | 2015 | VLDB |
| 9 | 3,110 | Similarity Evaluation on Tree-structured Data | 2005 | SIGMOD |
| 10 | 7,705 | An Incrementally Maintainable Index for Approximate Lookups in Hierarchical Data | 2006 | VLDB |