Back to papers
Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA
Summary: RAG-Sketch combines schema-driven query decomposition with skyline retrieval to select Pareto-optimal chunks balancing similarity and key-information coverage. A greedy set-cover merger removes redundant context access, improving long-context QA accuracy 24.5% while cutting token cost 58.5%.
(summarized by gpt-5.6-luna on Jul 26 2026)
Paper ID
7484
Venue
SIGMOD
Year
2026
Pagerank
5.093636e-05
Overall Rank
10,290 | 29.41%
DOI
10.1145/3802111
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
1.
Xinyi Zhu
(Hong Kong University of Science and Technology)
2.
Haoyang Li
(Hong Kong Polytechnic University)
3.
Yongqi Zhang
(Hong Kong University of Science and Technology)
4.
Lei Chen
(Hong Kong University of Science and Technology)
BibTeX Citation
Copy BibTeX
@inproceedings{zhu_sigmod26,
title = {{Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA}},
author = {Zhu, Xinyi and Li, Haoyang and Zhang, Yongqi and Chen, Lei},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802111},
url = {https://dl.acm.org/doi/10.1145/3802111},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
713
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
2024
VLDB
0.00014672521
2,816
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
2025
SIGMOD
8.0959781e-05
4,045
Logical and Physical Optimizations for SQL Query Execution over Large Language Models
2025
SIGMOD
6.9394654e-05
4,886
Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs
2024
VLDB
6.4614955e-05
6,101
LLM for Data Management
2024
VLDB
5.9774273e-05
6,254
In-depth Analysis of Graph-based RAG in a Unified Framework
2025
VLDB
5.9405174e-05
7,826
AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
2025
SIGMOD
5.53654e-05
7,871
TSGAssist: An Interactive Assistant Harnessing LLMs and RAG for Time Series Generation Recommendations and Benchmarking
2024
VLDB
5.5271792e-05
8,343
An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models
2024
VLDB
5.448279e-05
9,306
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
2025
VLDB
5.289545e-05
9,310
LLM-PBE: Assessing Data Privacy in Large Language Models
2024
VLDB
5.289545e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
10,382
LLM-Powered Interactive Graph Search: A Scalable and Practical Approach
2026
SIGMOD
2
6,765
AquaPipe: A Quality-Aware Pipeline for Knowledge Retrieval and Large Language Models
2025
SIGMOD
3
8,343
An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models
2024
VLDB
4
10,437
Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
2026
SIGMOD
5
10,405
AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA
2026
SIGMOD
6
10,499
SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQL
2026
SIGMOD
7
10,663
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
2025
SIGMOD
8
2,816
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
2025
SIGMOD
9
10,459
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
2026
SIGMOD
10
6,254
In-depth Analysis of Graph-based RAG in a Unified Framework
2025
VLDB