DBScholar

Back to papers

TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data

Summary: Proposes TASTI, a trainable semantic index replacing per-query proxies with embeddings so similar records share outputs. Theoretically ties embedding error to accuracy; empirically on five multimodal datasets, it builds 10x cheaper indexes and 24x faster proxy queries. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb355dedcd8d44137
Venue
SIGMOD
Year
2022
Pagerank
7.1341771e-05
Overall Rank
3,650 | 75.47%
DOI
10.1145/3514221.3517897

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{kang_sigmod22,
        title = {{TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data}},
        author = {Kang, Daniel and Guibas, John and Bailis, Peter D. and Hashimoto, Tatsunori and Zaharia, Matei},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517897},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517897},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
2,450 ThalamusDB: Approximate Query Processing on Multi-Modal Data 2024 SIGMOD 8.4474092e-05
3,778 Optimizing Video Analytics with Declarative Model Relationships 2023 VLDB 7.0259119e-05
4,557 Databases Unbound: Querying All of the World’s Bytes with AI 2024 VLDB 6.5356905e-05
5,551 Aero: Adaptive Query Processing of ML Queries 2025 SIGMOD 6.0872198e-05
6,583 Extract-Transform-Load for Video Streams 2023 VLDB 5.7438181e-05
7,447 EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User Interactions 2023 VLDB 5.5240718e-05
7,801 Accelerating Aggregation Queries on Unstructured Streams of Data 2023 VLDB 5.4507311e-05
8,997 Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees 2026 SIGMOD 5.2410834e-05
9,645 Self-Enhancing Video Data Management System for Compositional Events with Large Language Models 2025 SIGMOD 5.1453267e-05
10,107 TVM: A Tile-based Video Management Framework 2024 VLDB 5.0789354e-05
10,690 Task Cascades for Efficient Unstructured Data Processing 2026 SIGMOD 4.9793485e-05
10,845 Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees 2026 VLDB 4.9793485e-05
11,111 MAST: Towards Efficient Analytical Query Processing on Point Cloud Data 2025 SIGMOD 4.9793485e-05
11,210 Scalable Complex Event Processing on Video Streams 2025 SIGMOD 4.9793485e-05
11,509 Predictive and Near-Optimal Sampling for View Materialization in Video Databases 2024 SIGMOD 4.9793485e-05
11,596 Optimizing Video Queries with Declarative Clues 2024 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers