Chapter 22: Decision Trees
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 12 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
22.1 Tree-Based Learning
Tree-Based Learning (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Tree-Based Learning to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Tree-Based Learning
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Tree-Based Learning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.2 Root Nodes
Root Nodes (an entity or item in a graph). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Root Nodes to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Root Nodes
const graph = { A:['B','C'], B:['D'], C:['D'], D:[] };
const visited = new Set();
const queue = ['A'];
while(queue.length){
const node = queue.shift();
if(visited.has(node)) continue;
visited.add(node);
queue.push(...graph[node]);
}
console.log([...visited]);Code explanation
- The object stores a small graph as a list of neighbors for each node.
- A queue starts from node A and explores connected nodes breadth-first.
- The `visited` set prevents repeated work when different paths reach the same node.
- This traversal pattern is a foundation for graph features, connectivity checks, and many graph-learning workflows.
Expected result: The reachable nodes are printed in traversal order.
Practice exercise
Create a small real-world example for Root Nodes. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.3 Decision Nodes
Decision Nodes (an entity or item in a graph). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Decision Nodes to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Decision Nodes
const graph = { A:['B','C'], B:['D'], C:['D'], D:[] };
const visited = new Set();
const queue = ['A'];
while(queue.length){
const node = queue.shift();
if(visited.has(node)) continue;
visited.add(node);
queue.push(...graph[node]);
}
console.log([...visited]);Code explanation
- The object stores a small graph as a list of neighbors for each node.
- A queue starts from node A and explores connected nodes breadth-first.
- The `visited` set prevents repeated work when different paths reach the same node.
- This traversal pattern is a foundation for graph features, connectivity checks, and many graph-learning workflows.
Expected result: The reachable nodes are printed in traversal order.
Practice exercise
Create a small real-world example for Decision Nodes. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.4 Leaf Nodes
Leaf Nodes (an entity or item in a graph). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Leaf Nodes to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Leaf Nodes
const graph = { A:['B','C'], B:['D'], C:['D'], D:[] };
const visited = new Set();
const queue = ['A'];
while(queue.length){
const node = queue.shift();
if(visited.has(node)) continue;
visited.add(node);
queue.push(...graph[node]);
}
console.log([...visited]);Code explanation
- The object stores a small graph as a list of neighbors for each node.
- A queue starts from node A and explores connected nodes breadth-first.
- The `visited` set prevents repeated work when different paths reach the same node.
- This traversal pattern is a foundation for graph features, connectivity checks, and many graph-learning workflows.
Expected result: The reachable nodes are printed in traversal order.
Practice exercise
Create a small real-world example for Leaf Nodes. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.5 Splitting Features
Splitting Features (an input value or measurable property given to a model). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Splitting Features to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Splitting Features
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Splitting Features. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.6 Gini Impurity
Gini Impurity (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Gini Impurity to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Gini Impurity
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Gini Impurity. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.7 Entropy
Entropy (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Entropy to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Entropy
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Entropy. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.8 Information Gain
Information Gain (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Information Gain to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Information Gain
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Information Gain. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.9 Tree Depth
Tree Depth (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Tree Depth to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Tree Depth
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Tree Depth. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.10 Pruning
Pruning (removing less-important parameters or connections from a model). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Pruning to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Pruning
const model = input => input.reduce((a,b)=>a+b,0) / input.length;
const cache = new Map();
function predict(input){
const key = JSON.stringify(input);
if(cache.has(key)) return { value: cache.get(key), cached: true };
const value = model(input); cache.set(key,value);
return { value, cached: false };
}
console.log(predict([2,4,6]));
console.log(predict([2,4,6]));Code explanation
- `model()` stands in for a trained prediction function.
- `predict()` creates a stable key from the request so repeated inputs can be recognized.
- The first request computes and stores the result; the second request reuses it.
- This demonstrates a production concern—serving predictions efficiently—without depending on any particular deployment vendor.
Expected result: The first result is uncached and the second is returned from the cache.
Practice exercise
Create a small real-world example for Pruning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
22.11 Classification Trees
Classification Trees (predicting a category or class). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
An email filter receives a new message and decides whether it belongs to the 'spam' class or the 'not spam' class.
Coding example
// Classification Trees
const sigmoid = z => 1 / (1 + Math.exp(-z));
const weights = [0.8, -0.4];
const features = [2, 1];
const bias = -0.2;
const score = weights.reduce((sum, w, i) => sum + w * features[i], bias);
const probability = sigmoid(score);
const predictedClass = probability >= 0.5 ? 1 : 0;
console.log({ probability: probability.toFixed(3), predictedClass });Code explanation
- `weights`, `features`, and `bias` create a simple linear score.
- The sigmoid function converts any score into a value between 0 and 1.
- A threshold of 0.5 turns the probability into a class label.
- Printing both values helps you distinguish a model score from the final classification decision.
Expected result: A probability and a predicted class are printed.
Practice exercise
Create a second example for Classification Trees. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
22.12 Regression Trees
Regression Trees (predicting a continuous numerical value). Within Chapter 22, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Regression Trees to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Regression Trees
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Regression Trees. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 22 Review Questions and Answers
Q1. What is Tree-Based Learning?
Answer: Tree-Based Learning is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Root Nodes?
Answer: Root Nodes is an entity or item in a graph. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Decision Nodes?
Answer: Decision Nodes is an entity or item in a graph. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Leaf Nodes?
Answer: Leaf Nodes is an entity or item in a graph. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Splitting Features?
Answer: Splitting Features is an input value or measurable property given to a model. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Gini Impurity?
Answer: Gini Impurity is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Entropy?
Answer: Entropy is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Information Gain?
Answer: Information Gain is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Tree Depth?
Answer: Tree Depth is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Pruning?
Answer: Pruning is removing less-important parameters or connections from a model. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q11. What is Classification Trees?
Answer: Classification Trees is predicting a category or class. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q12. What is Regression Trees?
Answer: Regression Trees is predicting a continuous numerical value. In this chapter, focus on the input, the method or decision, and the result that should be checked.