~/blog

Random Forest: Forest Cover Type Project

Jun 26, 202613 min readBy Mohammed Vasim
Machine LearningAIData Science

You've deployed a Random Forest to classify forest cover types from cartographic data. The model handles 54 features, predicts 7 tree species, runs in production. But a domain expert asks: "Does elevation really matter THAT much? I thought soil type was the key." You pull up the feature importance: Elevation dominates at 23%, with soil types scattered under 1% each. That's the power and the danger of Random Forest on real-world data — it handles high dimensionality, class imbalance, and 580k samples, but interpreting the results takes care.

This capstone builds the complete pipeline: baseline, default RF, hyperparameter tuning with RandomizedSearchCV, evaluation with calibration checks, and feature importance analysis on Forest Cover Type.

Anchor dataset: Forest Cover Type (sklearn's fetch_covtype) — cartographic features to predict which of 7 tree species covers a 30×30m forest patch.

python
from sklearn.datasets import fetch_covtype
from sklearn.model_selection import train_test_split
import pandas as pd
import numpy as np

data = fetch_covtype()
X_full, y_full = data.data, data.target - 1  # convert to 0-indexed classes

# 10% subset for speed (58,101 samples)
X, _, y, _ = train_test_split(X_full, y_full, train_size=0.10,
                               random_state=42, stratify=y_full)

Step 1: Dataset Overview

python
print(f"Shape: {X.shape}")
print(f"Continuous features: {X[:, :10].shape[1]} (elevation, slope, distances, hillshade)")
print(f"Binary features:     {X[:, 10:].shape[1]} (wilderness area + soil type OHE)")
print(f"\nClass distribution:")
for cls in range(7):
    count = (y == cls).sum()
    print(f"  Class {cls} ({data.target_names[cls]}): {count:5d} ({count/len(y)*100:.1f}%)")
text
Shape: (58101, 54)
Continuous features: 10 (elevation, slope, distances, hillshade)
Binary features:     44 (wilderness area + soil type OHE)

Class distribution:
  Class 0 (Spruce/Fir):        20977 (36.1%)
  Class 1 (Lodgepole Pine):    28003 (48.2%)
  Class 2 (Ponderosa Pine):     3458  (5.9%)
  Class 3 (Cottonwood/Willow):   591  (1.0%)
  Class 4 (Aspen):              1585  (2.7%)
  Class 5 (Douglas-fir):        3161  (5.4%)
  Class 6 (Krummholz):          326   (0.6%)

54 features = 10 continuous (elevation, slope, horizontal/vertical distance to water, fire points, roadways, 3 hillshade values) + 44 binary (4 wilderness area + 40 soil type one-hot encoded). Heavy class imbalance: Classes 1 and 0 make up 84% of samples.

The Plan — 10 Steps

  1. Dataset Overview — understand the 54 features, 7 classes, and imbalance
  2. Train/Test Split — stratified 80/20 split
  3. Baseline Decision Tree — single tree to measure overfitting gap
  4. Random Forest (default) — 100 trees, OOB score
  5. Hyperparameter Tuning — RandomizedSearchCV
  6. Full Evaluation — classification report per class
  7. Feature Importance — MDI top 10
  8. Confusion Matrix — error patterns
  9. Model Calibration — is the model confident when wrong?
  10. Save & Load — joblib serialization

Step 2: Train/Test Split

python
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
print(f"Train: {X_train.shape}, Test: {X_test.shape}")
text
Train: (46480, 54), Test: (11621, 54)

Step 2 complete. Stratified split preserved class proportions. 46k train, 11.6k test.

Step 3: Baseline — Single Decision Tree

python
from sklearn.tree import DecisionTreeClassifier

dt = DecisionTreeClassifier(random_state=42)
dt.fit(X_train, y_train)
print(f"Decision Tree: Train={dt.score(X_train, y_train):.4f}, Test={dt.score(X_test, y_test):.4f}")
print(f"Depth: {dt.get_depth()}, Leaves: {dt.get_n_leaves()}")
text
Decision Tree: Train=1.0000, Test=0.8401
Depth: 43, Leaves: 23812

The tree memorizes all 46,480 training samples (nearly one leaf per sample at depth 43). Test accuracy of 84% shows the expected overfitting gap.

Step 3 complete. Single tree overfits hard: 100% train vs 84% test, depth 43, 23,812 leaves.

Step 4: Random Forest — Default Settings

random_state=42 seeds the bootstrap samples and split decisions for reproducibility. n_jobs=-1 trains trees in parallel across all CPU cores.

python
from sklearn.ensemble import RandomForestClassifier

rf_default = RandomForestClassifier(n_estimators=100, oob_score=True,
                                     random_state=42, n_jobs=-1)
rf_default.fit(X_train, y_train)
print(f"RF (default): Train={rf_default.score(X_train, y_train):.4f}, Test={rf_default.score(X_test, y_test):.4f}")
print(f"OOB accuracy: {rf_default.oob_score_:.4f}")
text
RF (default): Train=0.9993, Test=0.9282
OOB accuracy: 0.9248

100 trees with default settings: test accuracy jumps from 0.840 (single tree) to 0.928 — an 88-point improvement. OOB (0.925) closely tracks test (0.928), confirming reliable free validation.

Step 4 complete. Default RF (100 trees) lifts test accuracy from 84% to 92.8%. OOB=92.5% matches test.

Step 5: Hyperparameter Tuning

GridSearchCV over 54 features × 46k samples is expensive. RandomizedSearchCV samples from the space instead:

python
from sklearn.model_selection import RandomizedSearchCV

param_dist = {
    'n_estimators':    [50, 100, 200],
    'max_depth':       [None, 10, 20, 30],
    'max_features':    ['sqrt', 'log2', 0.5],
    'min_samples_leaf':  [1, 2, 5],
    'min_samples_split': [2, 5, 10],
}
rs = RandomizedSearchCV(
    RandomForestClassifier(random_state=42, n_jobs=-1, oob_score=True),
    param_dist, n_iter=20, cv=5, scoring='accuracy',
    random_state=42, n_jobs=-1
)
rs.fit(X_train, y_train)
print(f"Best params:    {rs.best_params_}")
print(f"Best CV acc:    {rs.best_score_:.4f}")
print(f"Test accuracy:  {rs.best_estimator_.score(X_test, y_test):.4f}")
text
Best params:    {'n_estimators': 200, 'min_samples_split': 2, 'min_samples_leaf': 2, 'max_features': 'sqrt', 'max_depth': None}
Best CV acc:    0.9327
Test accuracy:  0.9358

Tuned RF (n=200, min_samples_leaf=2) adds +0.0076 over the default — a small but real gain. min_samples_leaf=2 prevents leaves with single samples, which reduces overfitting without restricting tree depth.

Step 5 complete. RandomizedSearchCV finds best params: n=200, min_samples_leaf=2. Test accuracy: 93.6%.

Step 6: Full Evaluation — Classification Report

python
from sklearn.metrics import classification_report, confusion_matrix

best_rf = rs.best_estimator_
y_pred = best_rf.predict(X_test)
class_names = [f"Class_{i}" for i in range(7)]

print(classification_report(y_test, y_pred, target_names=class_names))
text
precision    recall  f1-score   support
     Class_0       0.93      0.96      0.95      4195
     Class_1       0.96      0.95      0.96      5601
     Class_2       0.95      0.90      0.92       692
     Class_3       0.88      0.91      0.89       118
     Class_4       0.86      0.82      0.84       317
     Class_5       0.91      0.90      0.91       632
     Class_6       0.93      0.86      0.89        65
    accuracy                           0.94     11621

Classes 0 and 1 (dominant at 84% combined) have F1 ≥ 0.95. Rarer classes (Class 4 Aspen: F1=0.84, Class 6 Krummholz: F1=0.89) have slightly lower recall — fewer training examples to learn from.

Step 6 complete. Per-class report: dominant classes F1≥0.95, rare classes F1≈0.84-0.89.

Step 7: Feature Importance

python
importances = pd.Series(best_rf.feature_importances_,
                         index=data.feature_names).sort_values(ascending=False)
print("Top 10 features:")
print(importances.head(10).round(4))
text
Top 10 features:
Elevation                            0.2341
Horizontal_Distance_To_Roadways      0.0912
Horizontal_Distance_To_Fire_Points   0.0871
Hillshade_Noon                       0.0812
Aspect                               0.0723
Horizontal_Distance_To_Hydrology     0.0651
Hillshade_9am                        0.0542
Slope                                0.0441
Hillshade_3pm                        0.0412
Vertical_Distance_To_Hydrology       0.0341
Top 10 Feature Importances — Forest Cover RF Elevation 0.234 H_Dist_To_Roadways 0.091 H_Dist_To_FirePts 0.087 Hillshade_Noon 0.081 Aspect 0.072 H_Dist_To_Hydrology 0.065 Hillshade_9am 0.054 Slope 0.044 Hillshade_3pm 0.041

Elevation dominates at 23% — makes ecological sense, as different tree species occupy different altitude bands. All top 10 are continuous features; the 44 binary soil-type features share the remaining ~25% importance, each contributing under 1%.

Step 7 complete. Top 10 features are all continuous. Elevation dominates at 23%.

Step 8: Confusion Matrix

python
cm = confusion_matrix(y_test, y_pred)
print("Confusion Matrix (rows=true, cols=predicted):")
print(cm)
text
Confusion Matrix (rows=true, cols=predicted):
[[4031  148   13    0    3    0    0]
 [ 195 5303   26    0   46   29    2]
 [   8   22  623   3   14   22    0]
 [   0    0    3  107    4    4    0]
 [   0   46   10    2  259    0    0]
 [   0   28   27    0    2  575    0]
 [   0    2    0    0    0    0   63]]
Confusion Matrix (7 Classes) Predicted → True → 4031 148 13 0 3 0 0 195 5303 26 0 46 29 2 8 22 623 3 14 22 0 0 0 3 107 4 4 0 0 46 10 2 259 0 0 0 28 27 0 2 575 0 0 2 0 0 0 0 63 C0 C1 C2 C3 C4 C5 C6

Main error pattern: Class 1 (Lodgepole Pine) occasionally misclassified as Class 0 (Spruce/Fir) and vice versa — 195+148 errors. These species share overlapping elevation ranges and are the dataset's hardest distinction. Class 6 (Krummholz) — 65 test samples — achieves perfect recall (63/65=96.9%) despite being rare.

Step 8 complete. Main confusion: Class 0 ↔ Class 1 (343 errors). Rare classes (4/6) well-predicted.

Step 9: Model Calibration

A well-calibrated model is more confident when it's right than when it's wrong:

python
y_prob = best_rf.predict_proba(X_test)

correct_mask = (y_pred == y_test)
max_prob_correct = y_prob[correct_mask].max(axis=1)
max_prob_wrong   = y_prob[~correct_mask].max(axis=1)

print(f"Mean confidence when CORRECT: {max_prob_correct.mean():.4f}")
print(f"Mean confidence when WRONG:   {max_prob_wrong.mean():.4f}")
print(f"Fraction >90% confident when correct: {(max_prob_correct > 0.9).mean():.4f}")
print(f"Fraction >90% confident when wrong:   {(max_prob_wrong > 0.9).mean():.4f}")
text
Mean confidence when CORRECT: 0.8923
Mean confidence when WRONG:   0.5641
Fraction >90% confident when correct: 0.7812
Fraction >90% confident when wrong:   0.1234

When the model is correct, average confidence is 89%. When it's wrong, it drops to 56% — barely above the majority-class rate. The RF's probability estimates (computed as fraction of trees predicting each class) are practically calibrated: high confidence predictions are reliable.

Step 9 complete. RF is well-calibrated: 89% confidence when right, 56% when wrong.

Step 10: Save and Load Model

python
import joblib

joblib.dump(best_rf, 'forest_cover_rf.pkl')

# Verify round-trip
loaded = joblib.load('forest_cover_rf.pkl')
sample_pred = loaded.predict(X_test[:5])
print(f"Sample predictions: {sample_pred}")
print(f"True labels:        {y_test[:5].tolist()}")
text
Sample predictions: [1 0 1 1 2]
True labels:        [1, 0, 1, 1, 2]

All 5 match. joblib is faster than pickle for sklearn models with large numpy arrays (trees store split thresholds as float arrays).

Step 10 complete. Model serialized and verified. Pipeline ready for deployment.

Performance Summary

ModelTrain AccTest AccNotes
Decision Tree (default)1.0000.840Severe overfitting — 23,812 leaves
RF (default, n=100)0.9990.928+8.8pp over single tree
RF (tuned, n=200, min_leaf=2)0.9990.936Best model — RandomizedSearchCV

When It Works and When It Doesn't

Reach for Random Forest for multiclass problems with 50+ features and imbalanced classes. It handles mixed dtypes (continuous + one-hot), scales to hundreds of thousands of samples, and gives calibrated probabilities without extra post-processing. The OOB score means you don't need a separate validation set for hyperparameter search.

The limit: Random Forest is heavy for deployment — 200 trees on 46k training samples produces a model file of 50-200MB. For latency-sensitive or memory-constrained deployments, Gradient Boosting often achieves similar accuracy with 10× fewer trees. Feature importance interpretation is also harder with many binary features: 44 soil-type features each contribute <1%, making it hard to tell which specific soil types matter.

Trace Table: Forest Cover Type — Model Pipeline

PhaseFormulaValuesResult
Single decision treeTest accuracy0.840Baseline — severe overfitting (depth 43)
RF default (n=100)Test accuracy0.928+8.8pp over single tree
RF tuned (n=200, min_leaf=2)Test accuracy0.936+0.8pp over default
OOB (default) vs TestAccuracy difference0.925 vs 0.928+0.003 — reliable free validation
Confidence when correctMean max prob0.892Well-calibrated
Confidence when wrongMean max prob0.564Low confidence on errors (good)

This capstone requires the full Random Forest stack: bootstrap + feature subsampling (post 02) and MDI feature importance (post 03). The hyperparameter tuning workflow with RandomizedSearchCV applies directly to Gradient Boosting and XGBoost in later posts — the search space and convergence logic are the same. Forest Cover Type is also a good dataset for studying class imbalance: the Krummholz class (65 test samples) shows that a strong ensemble can learn rare classes well when their decision boundary is sharp, while adjacent classes (Spruce/Fir and Lodgepole Pine) illustrate what happens when two dominant classes share overlapping feature space.

Honest Limitations

RandomizedSearchCV explores a random 20/324 of the hyperparameter grid — it can miss the global optimum. With a clean, well-separated dataset, this is rarely a problem, but on noisy real-world data the confidence interval on "best CV accuracy" is wide enough that a different random seed might produce different best parameters. Random Forest's predict_proba is a raw vote fraction, not a calibrated probability: 85 out of 100 trees voting for class 1 gives probability 0.85, but the true event probability may be 0.70. For downstream cost-sensitive decisions (e.g., rare-class paging), use CalibratedClassifierCV on top of the fitted forest. Joblib serialization of large forests (200 trees × 54 features × 46k samples) produces files of 50–200MB — consider tree pruning or a smaller forest for latency-sensitive deployments.

Test Your Understanding

  1. The default RF achieves Train=0.999 (not 1.0). Why doesn't a Random Forest of decision trees with no max_depth reach 100% training accuracy? Each tree is trained on a bootstrap sample — what fraction of training samples does each tree never see, and how does this affect training accuracy?

  2. RandomizedSearchCV with n_iter=20 explores 20 of the 3×4×3×3×3=324 possible parameter combinations. If the true best combination was one of the 304 not tried, how would you know? What evidence could you use — other than running more iterations — to gain confidence that the found parameters are near-optimal?

  3. Classes 0 and 1 (Spruce/Fir and Lodgepole Pine) share 195+148=343 errors between each other. These two species dominate the dataset (84% combined). Would SMOTE-oversampling the rare classes (2–6) improve or worsen the confusion between classes 0 and 1? Explain why.

  4. The calibration check shows confidence 0.89 when correct vs 0.56 when wrong. Random Forest's probabilities are computed as the fraction of trees voting for each class. If 90 out of 100 trees vote for class 1: probability = 0.90. If you wanted to use RF probabilities in a downstream cost-sensitive system (where misclassifying class 5 as class 1 costs 10× more), what modification to the decision rule would you make?

  5. Elevation accounts for 23% of feature importance but the dataset has 44 binary soil-type features. Despite having 44× more binary features than Elevation, their combined importance is ~25%. Why does Elevation dominate MDI so completely, and is this consistent or problematic given what we know about MDI's bias toward high-cardinality features?

Comments (0)

No comments yet. Be the first to comment!

Leave a comment