SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
Summary: SeDA introduces semantic document alignment for similar-passage search: semantic set-similarity over k-width windows, closing the gap between fast syntactic candidate generation and precise semantic matching. A candidate generator plus bound cascade exploits window overlap to prune >99% of expensive comparisons, yielding SBERT-quality F1 with near-syntactic runtimes. (summarized by gpt-5.4-mini on Apr 12 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Pranay Mundra (Yale University)
- 2. Daniel Kocher (University of Salzburg)
- 3. Martin Schäler (University of Salzburg)
- 4. Nikolaus Augsten (University of Salzburg)
BibTeX Citation
@article{mundra_vldb26,
title = {{SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora}},
author = {Mundra, Pranay and Kocher, Daniel and Schäler, Martin and Augsten, Nikolaus},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {6},
pages = {1332--1344},
doi = {10.14778/3797919.3797938},
url = {https://doi.org/10.14778/3797919.3797938},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,501 | An Empirical Evaluation of Set Similarity Join Techniques | 2016 | VLDB | 8.4975661e-05 |
| 3,040 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB | 7.8262287e-05 |
| 3,724 | Overlap Set Similarity Joins with Theoretical Guarantees | 2018 | SIGMOD | 7.1715735e-05 |
| 4,707 | SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints | 2017 | VLDB | 6.552423e-05 |
| 7,676 | Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts | 2021 | SIGMOD | 5.5686107e-05 |
| 7,745 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD | 5.5534781e-05 |
| 8,522 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD | 5.4119882e-05 |
| 11,447 | A Two-Level Signature Scheme for Stable Set Similarity Joins | 2023 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 12,282 | Aggregating Semantic Annotators | 2013 | VLDB |
| 2 | 13,314 | SemExplorer: A User Interface for Semantic Approach to Customized Dataset Search | 2025 | SIGMOD |
| 3 | 10,286 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD |
| 4 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 5 | 9,149 | Semantic SPARQL Similarity Search Over RDF Knowledge Graphs | 2016 | VLDB |
| 6 | 11,961 | Scalable Semantic Querying of Text | 2018 | VLDB |
| 7 | 10,206 | Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data | 2026 | SIGMOD |
| 8 | 9,383 | S^3AND: Efficient Subgraph Similarity Search Under Aggregated Neighbor Difference Semantics | 2025 | VLDB |
| 9 | 12,586 | SEDA: A System for Search, Exploration, Discovery, and Analysis of XML Data | 2008 | VLDB |
| 10 | 10,524 | A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs | 2026 | VLDB |