Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System
Summary: Pneuma is an end-to-end RAG system using LLMs to represent and retrieve tabular data, preserving schema and row context for accurate discovery. Evaluated on six real-world datasets, it outperforms full-text search and state-of-the-art RAG in accuracy and efficiency. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Muhammad Imam Luthfi Balaka (University of Indonesia)
- 2. David Alexander (University of Indonesia)
- 3. Qiming Wang (University of Chicago)
- 4. Yue Gong (University of Chicago)
- 5. Adila Krisnadhi (University of Indonesia)
- 6. Raul Castro Fernandez (University of Chicago)
BibTeX Citation
@inproceedings{balaka_sigmod25,
title = {{Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System}},
author = {Balaka, Muhammad Imam Luthfi and Alexander, David and Wang, Qiming and Gong, Yue and Krisnadhi, Adila and Fernandez, Raul Castro},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3725337},
url = {https://dl.acm.org/doi/10.1145/3725337},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,138 | The Pneuma Project: Reifying Information Needs as Relational Schemas to Automate Discovery, Guide Preparation, and Align Data with Intent | 2026 | CIDR | 5.093636e-05 |
| 10,431 | AutoDDG: Automated Dataset Description Generation using Large Language Models | 2026 | SIGMOD | 5.093636e-05 |
| 10,504 | Task Cascades for Efficient Unstructured Data Processing | 2026 | SIGMOD | 5.093636e-05 |
| 10,572 | Relational Deep Dive: Error-Aware Queries Over Unstructured Data | 2026 | VLDB | 5.093636e-05 |
| 10,618 | ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines | 2026 | VLDB | 5.093636e-05 |
| 10,627 | Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,245 | Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing | 2025 | CIDR | 0.00011507415 |
| 1,447 | Auctus: A Dataset Search Engine for Data Discovery and Augmentation | 2021 | VLDB | 0.00010760327 |
| 1,902 | Ground: A Data Context Service | 2017 | CIDR | 9.506714e-05 |
| 1,981 | ReAcTable: Enhancing ReAct for Table Question Answering | 2024 | VLDB | 9.3579557e-05 |
| 2,809 | Text2SQL is Not Enough: Unifying AI and Databases with TAG | 2025 | CIDR | 8.0994951e-05 |
| 2,956 | The Design of an LLM-powered Unstructured Analytics System | 2025 | CIDR | 7.9203461e-05 |
| 4,735 | Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation | 2022 | SIGMOD | 6.5315782e-05 |
| 7,220 | Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised Approach | 2023 | SIGMOD | 5.6679948e-05 |
Previous
Page 1 / 1
Next