Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models
Summary: Chameleon: a disaggregated heterogeneous accelerator architecture pairing FPGA vector-search accelerators with GPU LLM inference and CPU coordinators to independently scale retrieval and inference. Prototype yields up to 2.16× latency reduction and 3.18× throughput speedup vs CPU–GPU baselines. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Wenqi Jiang (ETH Zurich)
- 2. Marco Zeller (ETH Zurich)
- 3. Roger Waleffe (University of Wisconsin)
- 4. Torsten Hoefler (ETH Zurich)
- 5. Gustavo Alonso (ETH Zurich)
BibTeX Citation
@article{jiang_vldb25,
title = {{Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Models}},
author = {Jiang, Wenqi and Zeller, Marco and Waleffe, Roger and Hoefler, Torsten and Alonso, Gustavo},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {1},
pages = {42--52},
doi = {10.14778/3696435.3696439},
url = {https://doi.org/10.14778/3696435.3696439},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,254 | In-depth Analysis of Graph-based RAG in a Unified Framework | 2025 | VLDB | 5.9405174e-05 |
| 9,544 | SwiftSpatial: Spatial Joins on Modern Hardware | 2025 | SIGMOD | 5.2528121e-05 |
| 9,776 | Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal | 2025 | VLDB | 5.2209769e-05 |
| 10,209 | CMANNS: GPU-Accelerated Graph Index Construction for ANNS via Compute-Memory Disaggregation | 2026 | SIGMOD | 5.093636e-05 |
| 10,400 | Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search | 2026 | SIGMOD | 5.093636e-05 |
| 10,525 | Quantization Meets Projection: A Happy Marriage for Approximate k-Nearest Neighbor Search | 2026 | VLDB | 5.093636e-05 |
| 10,563 | Balancing the Blend: An Experimental Analysis of Trade-offs in Hybrid Search | 2026 | VLDB | 5.093636e-05 |
| 10,909 | LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 24 of 24 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,410 | TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware | 2026 | SIGMOD |
| 2 | 6,101 | LLM for Data Management | 2024 | VLDB |
| 3 | 13,290 | VecFlow-Chamfer: A GPU-based Data Management System for High-Performance Multi-Vector Search on Superchips | 2026 | SIGMOD |
| 4 | 8,343 | An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models | 2024 | VLDB |
| 5 | 13,343 | Database Perspective on LLM Inference Systems | 2025 | VLDB |
| 6 | 6,432 | RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference | 2026 | VLDB |
| 7 | 10,733 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 8 | 9,776 | Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal | 2025 | VLDB |
| 9 | 2,816 | Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation | 2025 | SIGMOD |
| 10 | 10,459 | From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation | 2026 | SIGMOD |