MoDora: Tree-Based Semi-Structured Document Analysis System
Summary: MoDora transforms OCR fragments into layout-aware components and a Component-Correlation Tree that preserves hierarchy, spatial distinctions, and cross-region links. Question-type-aware retrieval combines grid-based location search with LLM-guided semantic pruning, improving QA accuracy by 5.97–61.07%. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Bangrui Xu (Shanghai Jiao Tong University)
- 2. Qihang Yao (Shanghai Jiao Tong University)
- 3. Zirui Tang (Shanghai Jiao Tong University)
- 4. Xuanhe Zhou (Shanghai Jiao Tong University)
- 5. Yeye He (Microsoft)
- 6. Shihan Yu (Beihang University)
- 7. Qianqian Xu (Beihang University)
- 8. Bin Wang (Shanghai AI Laboratory)
- 9. Guoliang Li (Tsinghua University)
- 10. Conghui He (Shanghai AI Laboratory)
- 11. Fan Wu (Shanghai Jiao Tong University)
BibTeX Citation
@inproceedings{xu_sigmod26,
title = {{MoDora: Tree-Based Semi-Structured Document Analysis System}},
author = {Xu, Bangrui and Yao, Qihang and Tang, Zirui and Zhou, Xuanhe and He, Yeye and Yu, Shihan and Xu, Qianqian and Wang, Bin and Li, Guoliang and He, Conghui and Wu, Fan},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802089},
url = {https://dl.acm.org/doi/10.1145/3802089},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00014672521 |
| 1,245 | Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing | 2025 | CIDR | 0.00011507415 |
| 6,371 | QUEST: Query Optimization in Unstructured Document Analysis | 2025 | VLDB | 5.8962187e-05 |
| 10,403 | ST-Raptor: LLM-Powered Semi-Structured Table Question Answering | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,671 | Question Answering Over Knowledge Graphs: Question Understanding Via Template Decomposition | 2018 | VLDB |
| 2 | 4,886 | Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs | 2024 | VLDB |
| 3 | 10,286 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD |
| 4 | 10,719 | Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents | 2025 | SIGMOD |
| 5 | 1,343 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB |
| 6 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 7 | 13,339 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 8 | 8,343 | An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models | 2024 | VLDB |
| 9 | 8,340 | Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents | 2025 | VLDB |
| 10 | 10,403 | ST-Raptor: LLM-Powered Semi-Structured Table Question Answering | 2026 | SIGMOD |