Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees
Summary: Blocking-augmented sampling (BaS) combines embedding blocking with sampling for ML-based joins over unstructured data. It adaptively handles false negatives/positives, providing confidence intervals and up to 19× lower error than prior methods. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yuxuan Zhu (University of Illinois Urbana-Champaign)
- 2. Tengjun Jin (University of Illinois Urbana-Champaign)
- 3. Chenghao Mo (University of Illinois Urbana-Champaign)
- 4. Daniel Kang (University of Illinois Urbana-Champaign)
BibTeX Citation
@inproceedings{zhu_sigmod26,
title = {{Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees}},
author = {Zhu, Yuxuan and Jin, Tengjun and Mo, Chenghao and Kang, Daniel},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802004},
url = {https://dl.acm.org/doi/10.1145/3802004},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 27 of 27 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 2 | 2,120 | Comparative Analysis of Approximate Blocking Techniques for Entity Resolution | 2016 | VLDB |
| 3 | 136 | Join Synopses for Approximate Query Answering | 1999 | SIGMOD |
| 4 | 3,907 | Accelerating Approximate Aggregation Queries with Expensive Predicates | 2021 | VLDB |
| 5 | 10,169 | Towards Output-Optimal Uniform Sampling and Approximate Counting for Join-Project Queries | 2026 | PODS |
| 6 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 7 | 10,634 | Efficient Approximate Query Processing with Block Sampling | 2025 | CIDR |
| 8 | 7,048 | Learning to Sample: Counting with Complex Queries | 2020 | VLDB |
| 9 | 9,052 | Fast Approximate Similarity Join in Vector Databases | 2025 | SIGMOD |
| 10 | 7,335 | Reservoir Sampling over Joins | 2024 | SIGMOD |