SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
Summary: SeDA introduces semantic document alignment for similar-passage search: semantic set-similarity over k-width windows, closing the gap between fast syntactic candidate generation and precise semantic matching. A candidate generator plus bound cascade exploits window overlap to prune >99% of expensive comparisons, yielding SBERT-quality F1 with near-syntactic runtimes. (summarized by gpt-5.4-mini on Apr 12 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Pranay Mundra (Yale University)
- 2. Daniel Kocher (University of Salzburg)
- 3. Martin Schäler (University of Salzburg)
- 4. Nikolaus Augsten (University of Salzburg)
BibTeX Citation
@article{mundra_vldb26,
title = {{SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora}},
author = {Mundra, Pranay and Kocher, Daniel and Schäler, Martin and Augsten, Nikolaus},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {6},
pages = {1332--1344},
doi = {10.14778/3797919.3797938},
url = {https://doi.org/10.14778/3797919.3797938},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,513 | An Empirical Evaluation of Set Similarity Join Techniques | 2016 | VLDB | 8.3679178e-05 |
| 2,921 | Leveraging Set Relations in Exact Set Similarity Join | 2017 | VLDB | 7.8519256e-05 |
| 3,596 | Overlap Set Similarity Joins with Theoretical Guarantees | 2018 | SIGMOD | 7.1790375e-05 |
| 4,571 | SILKMOTH: An Efficient Method for Finding Related Sets with Maximum Matching Constraints | 2017 | VLDB | 6.5286445e-05 |
| 7,833 | Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts | 2021 | SIGMOD | 5.443666e-05 |
| 7,906 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD | 5.428873e-05 |
| 8,689 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD | 5.2905577e-05 |
| 11,759 | A Two-Level Signature Scheme for Stable Set Similarity Joins | 2023 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 13,636 | SemExplorer: A User Interface for Semantic Approach to Customized Dataset Search | 2025 | SIGMOD |
| 2 | 10,498 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD |
| 3 | 10,858 | Sema: A High-performance System for LLM-based Semantic Query Processing | 2026 | VLDB |
| 4 | 11,529 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 5 | 9,323 | Semantic SPARQL Similarity Search Over RDF Knowledge Graphs | 2016 | VLDB |
| 6 | 12,259 | Scalable Semantic Querying of Text | 2018 | VLDB |
| 7 | 10,422 | Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data | 2026 | SIGMOD |
| 8 | 9,568 | S^3AND: Efficient Subgraph Similarity Search Under Aggregated Neighbor Difference Semantics | 2025 | VLDB |
| 9 | 12,876 | SEDA: A System for Search, Exploration, Discovery, and Analysis of XML Data | 2008 | VLDB |
| 10 | 10,708 | A Semantics-aware Approach for Graph Edit Distance Estimation over Knowledge Graphs | 2026 | VLDB |