Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates
Summary: ReTree eliminates redundant decision-tree feature tests when ML inference appears in SQL predicates, targeting sibling- and ancestor-induced redundancy. Its subtree collapse and recombination techniques, implemented in DuckDB, deliver 2.56× average speedup. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Mingxi Liu (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 2. Zhengyuan Ding (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 3. Chenyang Zhang (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 4. Qingfeng Pan (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 5. Huayou Su (National University of Defense Technology)
- 6. Zhao Zhang (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 7. Chen Xu (East China Normal University; Engineering Research Center of Blockchain Data Management; Shanghai Engineering Research Center of Big Data Management)
- 8. Qingsong Ruan (China Electronics Technology Kingbase (Beijing) Technologies Inc.)
BibTeX Citation
@inproceedings{liu_sigmod26,
title = {{Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates}},
author = {Liu, Mingxi and Ding, Zhengyuan and Zhang, Chenyang and Pan, Qingfeng and Su, Huayou and Zhang, Zhao and Xu, Chen and Ruan, Qingsong},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802048},
url = {https://dl.acm.org/doi/10.1145/3802048},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 22 of 22 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,355 | QueryFormer: A Tree Transformer Model for Query Plan Representation | 2022 | VLDB |
| 2 | 4,114 | Optimizing Machine Learning Inference Queries with Correlative Proxy Models | 2022 | VLDB |
| 3 | 5,767 | A Comparative Study and Component Analysis of Query Plan Representation Techniques in ML4DB Studies | 2024 | VLDB |
| 4 | 1,994 | RainForest - A Framework for Fast Decision Tree Construction of Large Datasets | 1998 | VLDB |
| 5 | 465 | An End-to-End Learning-based Cost Estimator | 2020 | VLDB |
| 6 | 697 | Selectivity Estimation for Range Predicates using Lightweight Models | 2019 | VLDB |
| 7 | 6,585 | JoinBoost: Grow Trees Over Normalized Data Using Only SQL | 2023 | VLDB |
| 8 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 9 | 9,874 | Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests | 2023 | VLDB |
| 10 | 8,572 | T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision Trees | 2025 | SIGMOD |