Back to papers
Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA
Summary: RAG-Sketch combines schema-driven query decomposition with skyline retrieval to select Pareto-optimal chunks balancing similarity and key-information coverage. A greedy set-cover merger removes redundant context access, improving long-context QA accuracy 24.5% while cutting token cost 58.5%.
(summarized by gpt-5.6-luna on Jul 26 2026)
Paper ID
he4bacf954394752c
Venue
SIGMOD
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,502 | 29.40%
DOI
10.1145/3802111
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
1.
Xinyi Zhu
(Hong Kong University of Science and Technology)
2.
Haoyang Li
(Hong Kong Polytechnic University)
3.
Yongqi Zhang
(Hong Kong University of Science and Technology)
4.
Lei Chen
(Hong Kong University of Science and Technology)
BibTeX Citation
Copy BibTeX
@inproceedings{zhu_sigmod26,
title = {{Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA}},
author = {Zhu, Xinyi and Li, Haoyang and Zhang, Yongqi and Chen, Lei},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802111},
url = {https://dl.acm.org/doi/10.1145/3802111},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
501
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
2024
VLDB
0.00017267905
2,638
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
2025
SIGMOD
8.1839298e-05
3,466
In-depth Analysis of Graph-based RAG in a Unified Framework
2025
VLDB
7.2783993e-05
3,743
Logical and Physical Optimizations for SQL Query Execution over Large Language Models
2025
SIGMOD
7.0586112e-05
4,718
Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs
2024
VLDB
6.4551989e-05
5,149
LLM for Data Management
2024
VLDB
6.2551644e-05
6,473
AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
2025
SIGMOD
5.7739763e-05
7,318
An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models
2024
VLDB
5.5547304e-05
7,987
TSGAssist: An Interactive Assistant Harnessing LLMs and RAG for Time Series Generation Recommendations and Benchmarking
2024
VLDB
5.4119536e-05
9,484
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
2025
VLDB
5.1708619e-05
9,487
LLM-PBE: Assessing Data Privacy in Large Language Models
2024
VLDB
5.1708619e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
10,626
Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
2026
SIGMOD
2
10,598
AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA
2026
SIGMOD
3
10,686
SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQL
2026
SIGMOD
4
10,791
BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
2026
VLDB
5
9,475
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
2025
SIGMOD
6
10,890
RAGPerf: An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
2026
VLDB
7
2,638
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
2025
SIGMOD
8
10,646
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
2026
SIGMOD
9
3,466
In-depth Analysis of Graph-based RAG in a Unified Framework
2025
VLDB
10
10,781
QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented Generation
2026
VLDB