Back to papers
SemBench: A Benchmark for Semantic Query Processing Engines
Summary: SemBench benchmarks LLM-powered semantic query engines that extend SQL with natural-language operators over multimodal data. It spans diverse scenarios, modalities, and operators, exposing strengths and weaknesses across academic and industrial systems.
(summarized by gpt-5.6-luna on Jul 09 2026)
- Paper ID
- 14316
- Venue
- VLDB
- Year
- 2026
- Pagerank
- 4.1905499e-05
- Overall Rank
- 10,277 | 28.58%
- DOI
-
10.14778/3811243.3811249
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 71 |
How Good Are Query Optimizers, Really? |
2016 |
VLDB |
0.00059446482 |
| 94 |
CrowdDB: Answering Queries with Crowdsourcing |
2011 |
SIGMOD |
0.00051273089 |
| 167 |
The Snowflake Elastic Data Warehouse |
2016 |
SIGMOD |
0.00039408116 |
| 246 |
Crowdsourced Databases: Query Processing with People |
2011 |
CIDR |
0.00030952631 |
| 266 |
Human-powered Sorts and Joins |
2012 |
VLDB |
0.00029884758 |
| 997 |
CAESURA: Language Models as Multi-Modal Query Planners |
2024 |
CIDR |
0.00014726927 |
| 1,839 |
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing |
2025 |
VLDB |
0.00010351287 |
| 2,255 |
Counting with the Crowd |
2013 |
VLDB |
9.1846281e-05 |
| 5,149 |
Abacus: A Cost-Based Optimizer for Semantic Operator Systems |
2026 |
VLDB |
5.655398e-05 |
| 5,206 |
ThalamusDB: Approximate Query Processing on Multi-Modal Data |
2024 |
SIGMOD |
5.625641e-05 |
| 5,799 |
Learned Approximate Query Processing: Make it Light, Accurate and Fast |
2021 |
CIDR |
5.3219666e-05 |
| 8,414 |
PairwiseHist: Fast, Accurate and Space-Efficient Approximate Query Processing with Data Compression |
2024 |
VLDB |
4.5135713e-05 |
| 9,238 |
PilotDB: Database-Agnostic Online Approximate Query Processing with A Priori Error Guarantees |
2025 |
SIGMOD |
4.3648789e-05 |
| 9,242 |
Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB |
2025 |
VLDB |
4.3648789e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 9,210 |
SmartBench: Demonstrating Automatic Generation of Comprehensive Benchmarks for Question Answering Over Knowledge Graphs |
2022 |
VLDB |
4.3689343e-05 |
| 7,887 |
SQLStorm: Taking Database Benchmarking into the LLM Era |
2025 |
VLDB |
4.6218382e-05 |
| 13,123 |
SemExplorer: A User Interface for Semantic Approach to Customized Dataset Search |
2025 |
SIGMOD |
- |
| 10,144 |
Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization |
2026 |
SIGMOD |
4.1905499e-05 |
| 9,989 |
Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics |
2026 |
CIDR |
4.1905499e-05 |
| 5,363 |
An In-Depth Benchmarking of Text-to-SQL Systems |
2021 |
SIGMOD |
5.5467941e-05 |
| 9,992 |
Leveraging Query Optimizers to Verify the Soundness of LLM-based Query Rewrites for Real-World Workloads, and More! |
2026 |
CIDR |
4.1905499e-05 |
| 9,973 |
BenchPress: A Human-in-the-Loop Annotation System for Rapid Text-to-SQL Benchmark Curation |
2026 |
CIDR |
4.1905499e-05 |
| 10,221 |
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions |
2026 |
VLDB |
4.1905499e-05 |
| 8,464 |
Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS |
2025 |
VLDB |
4.5003888e-05 |