Chapter 26: Classification Evaluation
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 15 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
26.1 Confusion Matrix
Confusion Matrix (a table comparing predicted classes with actual classes). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
If a medical classifier predicts 100 cases, a confusion matrix separates correct positives, false positives, correct negatives, and false negatives so you can see exactly what kinds of mistakes occurred.
Coding example
// Confusion Matrix
const a = [2, 4, 6];
const b = [1, 3, 5];
const dot = a.reduce((sum, value, i) => sum + value * b[i], 0);
const magnitude = Math.sqrt(a.reduce((sum, value) => sum + value ** 2, 0));
console.log({ dot, magnitude: magnitude.toFixed(2) });Code explanation
- The arrays `a` and `b` represent small numeric vectors so the calculation stays easy to inspect.
- `reduce()` walks through the values and combines them into one result, which is useful for many linear-algebra operations.
- The magnitude calculation squares each value, adds the squares, and takes the square root.
- The final object prints values you can compare by hand before using the same idea with larger data.
Expected result: A dot-product value and a vector magnitude are printed.
Practice exercise
Create a second example for Confusion Matrix. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
26.2 True Positives
True Positives (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use True Positives to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// True Positives
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for True Positives. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.3 False Positives
False Positives (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use False Positives to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// False Positives
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for False Positives. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.4 True Negatives
True Negatives (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use True Negatives to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// True Negatives
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for True Negatives. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.5 False Negatives
False Negatives (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use False Negatives to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// False Negatives
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for False Negatives. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.6 Accuracy
Accuracy (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small real-world project where Accuracy is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Accuracy
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Accuracy. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
26.7 Precision
Precision (the share of predicted positives that are actually positive). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
If a system marks 10 emails as spam and 8 really are spam, its precision is 8 out of 10, or 80%.
Coding example
// Precision
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a second example for Precision. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
26.8 Recall
Recall (the share of actual positives that the model correctly finds). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
If there are 10 real spam emails and the system finds 8 of them, recall is 8 out of 10, or 80%.
Coding example
// Recall
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a second example for Recall. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
26.9 Specificity
Specificity (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Specificity to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Specificity
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Specificity. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.10 F1 Score
F1 Score (a balanced combination of precision and recall). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use F1 Score to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// F1 Score
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a small real-world example for F1 Score. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.11 ROC Curve
ROC Curve (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use ROC Curve to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// ROC Curve
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a small real-world example for ROC Curve. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.12 AUC
AUC (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use AUC to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// AUC
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a small real-world example for AUC. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.13 Precision-Recall Curves
Precision-Recall Curves (the share of predicted positives that are actually positive). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
If a system marks 10 emails as spam and 8 really are spam, its precision is 8 out of 10, or 80%.
Coding example
// Precision-Recall Curves
const truth = [1,1,0,1,0,0,1,0];
const pred = [1,0,0,1,1,0,1,0];
let tp=0,fp=0,fn=0,tn=0;
truth.forEach((y,i)=>{ const p=pred[i]; if(y===1&&p===1)tp++; else if(y===0&&p===1)fp++; else if(y===1&&p===0)fn++; else tn++; });
const precision = tp / (tp + fp);
const recall = tp / (tp + fn);
console.log({tp,fp,fn,tn,precision:precision.toFixed(2),recall:recall.toFixed(2)});Code explanation
- `truth` holds correct labels and `pred` holds model predictions in the same order.
- The loop counts true positives, false positives, false negatives, and true negatives.
- Precision asks how many predicted positives were correct, while recall asks how many real positives were found.
- These values reveal different kinds of classification errors that accuracy alone can hide.
Expected result: Confusion-matrix counts, precision, and recall are printed.
Practice exercise
Create a second example for Precision-Recall Curves. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
26.14 Multiclass Metrics
Multiclass Metrics (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Multiclass Metrics to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Multiclass Metrics
const values = [12, 15, 11, 18, 14, 16];
const mean = values.reduce((sum, x) => sum + x, 0) / values.length;
const variance = values.reduce((sum, x) => sum + (x - mean) ** 2, 0) / values.length;
const std = Math.sqrt(variance);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2) });Code explanation
- `values` is a tiny dataset that can be checked manually.
- The mean is the total divided by the number of observations.
- Variance measures average squared distance from the mean, and the square root of variance gives standard deviation.
- These summary values help you understand the scale and spread of data before choosing or evaluating a model.
Expected result: The mean and standard deviation are printed.
Practice exercise
Create a small real-world example for Multiclass Metrics. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
26.15 Threshold Selection
Threshold Selection (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 26, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Threshold Selection to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Threshold Selection
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Threshold Selection. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 26 Review Questions and Answers
Q1. What is Confusion Matrix?
Answer: Confusion Matrix is a table comparing predicted classes with actual classes. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is True Positives?
Answer: True Positives is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is False Positives?
Answer: False Positives is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is True Negatives?
Answer: True Negatives is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is False Negatives?
Answer: False Negatives is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Accuracy?
Answer: Accuracy is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Precision?
Answer: Precision is the share of predicted positives that are actually positive. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Recall?
Answer: Recall is the share of actual positives that the model correctly finds. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Specificity?
Answer: Specificity is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is F1 Score?
Answer: F1 Score is a balanced combination of precision and recall. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q11. What is ROC Curve?
Answer: ROC Curve is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q12. What is AUC?
Answer: AUC is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q13. What is Precision-Recall Curves?
Answer: Precision-Recall Curves is the share of predicted positives that are actually positive. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q14. What is Multiclass Metrics?
Answer: Multiclass Metrics is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q15. What is Threshold Selection?
Answer: Threshold Selection is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.