Chapter 23: Random Forests
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 10 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
23.1 Ensemble Learning
Ensemble Learning (a system that combines predictions from multiple models). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Ensemble Learning to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Ensemble Learning
const samples = [
{ value: 2, label: 0 }, { value: 4, label: 0 },
{ value: 7, label: 1 }, { value: 9, label: 1 }
];
const threshold = 5;
const predict = value => value <= threshold ? 0 : 1;
const correct = samples.filter(s => predict(s.value) === s.label).length;
console.log({ threshold, accuracy: correct / samples.length });Code explanation
- The examples contain one feature named `value` and a known class label.
- A threshold acts like one simple decision-tree split.
- The prediction function sends values to one of two branches based on that threshold.
- Counting correct predictions shows how a split can be evaluated before it is combined with more splits or more trees.
Expected result: The split threshold and its accuracy on the toy data are printed.
Practice exercise
Create a small real-world example for Ensemble Learning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.2 Bootstrap Sampling
Bootstrap Sampling (selecting a subset of a larger population or dataset). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Bootstrap Sampling to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Bootstrap Sampling
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Bootstrap Sampling. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.3 Bagging
Bagging (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Bagging to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Bagging
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Bagging. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.4 Random Feature Selection
Random Feature Selection (an input value or measurable property given to a model). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Random Feature Selection to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Random Feature Selection
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Random Feature Selection. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.5 Random Forest Classification
Random Forest Classification (predicting a category or class). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
An email filter receives a new message and decides whether it belongs to the 'spam' class or the 'not spam' class.
Coding example
// Random Forest Classification
const sigmoid = z => 1 / (1 + Math.exp(-z));
const weights = [0.8, -0.4];
const features = [2, 1];
const bias = -0.2;
const score = weights.reduce((sum, w, i) => sum + w * features[i], bias);
const probability = sigmoid(score);
const predictedClass = probability >= 0.5 ? 1 : 0;
console.log({ probability: probability.toFixed(3), predictedClass });Code explanation
- `weights`, `features`, and `bias` create a simple linear score.
- The sigmoid function converts any score into a value between 0 and 1.
- A threshold of 0.5 turns the probability into a class label.
- Printing both values helps you distinguish a model score from the final classification decision.
Expected result: A probability and a predicted class are printed.
Practice exercise
Create a second example for Random Forest Classification. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
23.6 Random Forest Regression
Random Forest Regression (an ensemble of many randomized decision trees whose predictions are combined). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
A model receives the size, age, and location of a house and predicts a numerical price such as $850,000.
Coding example
// Random Forest Regression
const samples = [
{ value: 2, label: 0 }, { value: 4, label: 0 },
{ value: 7, label: 1 }, { value: 9, label: 1 }
];
const threshold = 5;
const predict = value => value <= threshold ? 0 : 1;
const correct = samples.filter(s => predict(s.value) === s.label).length;
console.log({ threshold, accuracy: correct / samples.length });Code explanation
- The examples contain one feature named `value` and a known class label.
- A threshold acts like one simple decision-tree split.
- The prediction function sends values to one of two branches based on that threshold.
- Counting correct predictions shows how a split can be evaluated before it is combined with more splits or more trees.
Expected result: The split threshold and its accuracy on the toy data are printed.
Practice exercise
Create a second example for Random Forest Regression. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
23.7 Feature Importance
Feature Importance (an input value or measurable property given to a model). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Feature Importance to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Feature Importance
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Feature Importance. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.8 Out-of-Bag Evaluation
Out-of-Bag Evaluation (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Out-of-Bag Evaluation to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Out-of-Bag Evaluation
const values = [12, 15, 11, 18, 14, 16];
const mean = values.reduce((sum, x) => sum + x, 0) / values.length;
const variance = values.reduce((sum, x) => sum + (x - mean) ** 2, 0) / values.length;
const std = Math.sqrt(variance);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2) });Code explanation
- `values` is a tiny dataset that can be checked manually.
- The mean is the total divided by the number of observations.
- Variance measures average squared distance from the mean, and the square root of variance gives standard deviation.
- These summary values help you understand the scale and spread of data before choosing or evaluating a model.
Expected result: The mean and standard deviation are printed.
Practice exercise
Create a small real-world example for Out-of-Bag Evaluation. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.9 Hyperparameters
Hyperparameters (a model or training setting chosen outside the learned parameters). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Hyperparameters to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Hyperparameters
const choices = [0.01, 0.05, 0.1, 0.2];
const evaluate = value => 1 - Math.abs(value - 0.08);
const results = choices.map(value => ({ value, score: evaluate(value) }));
results.sort((a,b) => b.score - a.score);
console.log('best choice:', results[0]);Code explanation
- `choices` represents candidate settings that could be tried automatically.
- `evaluate()` stands in for a validation process that assigns each candidate a score.
- All candidates are evaluated and sorted from best to worst.
- The highest-scoring setting is selected, demonstrating the core search loop behind many tuning systems.
Expected result: The best candidate setting and its score are printed.
Practice exercise
Create a small real-world example for Hyperparameters. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
23.10 Advantages and Limitations
Advantages and Limitations (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 23, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Advantages and Limitations to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Advantages and Limitations
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Advantages and Limitations. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 23 Review Questions and Answers
Q1. What is Ensemble Learning?
Answer: Ensemble Learning is a system that combines predictions from multiple models. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Bootstrap Sampling?
Answer: Bootstrap Sampling is selecting a subset of a larger population or dataset. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Bagging?
Answer: Bagging is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Random Feature Selection?
Answer: Random Feature Selection is an input value or measurable property given to a model. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Random Forest Classification?
Answer: Random Forest Classification is predicting a category or class. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Random Forest Regression?
Answer: Random Forest Regression is an ensemble of many randomized decision trees whose predictions are combined. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Feature Importance?
Answer: Feature Importance is an input value or measurable property given to a model. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Out-of-Bag Evaluation?
Answer: Out-of-Bag Evaluation is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Hyperparameters?
Answer: Hyperparameters is a model or training setting chosen outside the learned parameters. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Advantages and Limitations?
Answer: Advantages and Limitations is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.