Self-Enhancing Video Data Management System for Compositional Events with Large Language Models
Summary: VOCAL-UDF enables compositional video queries without predefined modules via auto-generated UDFs in an LLM-based model. It supports program UDFs and distilled UDFs, generates candidates to resolve intent, and uses active learning to select the best. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Enhao Zhang (University of Washington)
- 2. Nicole Sullivan (University of Washington)
- 3. Brandon Haynes (Microsoft)
- 4. Ranjay Krishna (University of Washington)
- 5. Magdalena Balazinska (University of Washington)
BibTeX Citation
@inproceedings{zhang_sigmod25,
title = {{Self-Enhancing Video Data Management System for Compositional Events with Large Language Models}},
author = {Zhang, Enhao and Sullivan, Nicole and Haynes, Brandon and Krishna, Ranjay and Balazinska, Magdalena},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725352},
url = {https://dl.acm.org/doi/10.1145/3725352},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,286 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 25 of 25 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,518 | DeepVQL: Deep Video Queries on PostgreSQL | 2023 | VLDB |
| 2 | 4,429 | Evaluating Temporal Queries Over Video Feeds | 2021 | SIGMOD |
| 3 | 2,384 | DeepLens: Towards a Visual Data Management System | 2019 | CIDR |
| 4 | 9,408 | SketchQL: Video Moment Querying with a Visual Query Interface | 2024 | SIGMOD |
| 5 | 2,702 | Panorama: A Data System for Unbounded Vocabulary Querying over Video | 2020 | VLDB |
| 6 | 10,917 | Deja Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse | 2025 | VLDB |
| 7 | 11,483 | EQUI-VOCAL Demonstration: Synthesizing Video Queries from User Interactions | 2023 | VLDB |
| 8 | 11,268 | Optimizing Video Queries with Declarative Clues | 2024 | VLDB |
| 9 | 8,362 | EQUI-VOCAL: Synthesizing Queries for Compositional Video Events from Limited User Interactions | 2023 | VLDB |
| 10 | 5,966 | VOCAL: Video Organization and Interactive Compositional AnaLytics | 2022 | CIDR |