Back to papers
PLANET: Massively Parallel Learning of Tree Ensembles with MapReduce
Summary: PLANET leverages MapReduce to train tree ensembles on massive datasets using commodity hardware. It frames tree learning as distributed MapReduce steps, enabling scalable construction of classification/regression trees and ensembles on commodity clusters, demonstrated on computational advertising.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 9906
- Venue
- VLDB
- Year
- 2009
- Pagerank
- 8.401513e-05
- Overall Rank
- 2,636 | 81.69%
- DOI
-
-
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 10 of 10 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 1,165 |
Simulation of Database-Valued Markov Chains Using SimSQL |
2013 |
SIGMOD |
0.00013567206 |
| 1,899 |
VF2Boost: Very Fast Vertical Federated Gradient Boosting for Cross-Enterprise Learning |
2021 |
SIGMOD |
0.00010171063 |
| 1,942 |
SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging |
2021 |
SIGMOD |
0.00010010569 |
| 2,340 |
Efficient Processing of Data Warehousing Queries in a Split Execution Environment |
2011 |
SIGMOD |
9.001663e-05 |
| 2,660 |
Vertica-ML: Distributed Machine Learning in Vertica Database |
2020 |
SIGMOD |
8.3614044e-05 |
| 2,714 |
Minimal MapReduce Algorithms |
2013 |
SIGMOD |
8.2426646e-05 |
| 4,397 |
Smurf: Self-Service String Matching Using Random Forests |
2019 |
VLDB |
6.2142485e-05 |
| 7,293 |
Optimization for iterative queries on MapReduce |
2014 |
VLDB |
4.7668182e-05 |
| 8,050 |
Lowering the Latency of Data Processing Pipelines Through FPGA based Hardware Acceleration |
2020 |
VLDB |
4.5933331e-05 |
| 11,253 |
Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests |
2023 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 11,450 |
Grouped Learning: Group-By Model Selection Workloads |
2021 |
SIGMOD |
4.1905499e-05 |
| 4,601 |
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches |
2021 |
VLDB |
6.05274e-05 |
| 3,092 |
Scalable and Efficient Full-Graph GNN Training for Large Graphs |
2023 |
SIGMOD |
7.5869574e-05 |
| 8,973 |
Planting Trees for scalable and efficient Canonical Hub Labeling |
2020 |
VLDB |
4.4148296e-05 |
| 13,356 |
M3: Scaling Up Machine Learning via Memory Mapping |
2016 |
SIGMOD |
- |
| 12,738 |
Parallel Mining Algorithms for Generalized Association Rules with Classification Hierarchy |
1998 |
SIGMOD |
4.1905499e-05 |
| 1,407 |
Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML |
2014 |
VLDB |
0.00012163413 |
| 9,225 |
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning |
2021 |
VLDB |
4.3656789e-05 |
| 7,771 |
Mining Tree-Structured Data on Multicore Systems |
2009 |
VLDB |
4.6512834e-05 |
| 1,458 |
RainForest - A Framework for Fast Decision Tree Construction of Large Datasets |
1998 |
VLDB |
0.00011888676 |