Chapter 21: k-Nearest Neighbors
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 10 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
21.1 Instance-Based Learning
Instance-Based Learning (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Instance-Based Learning to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Instance-Based Learning
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Instance-Based Learning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.2 Distance Metrics
Distance Metrics (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Distance Metrics to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Distance Metrics
const values = [12, 15, 11, 18, 14, 16];
const mean = values.reduce((sum, x) => sum + x, 0) / values.length;
const variance = values.reduce((sum, x) => sum + (x - mean) ** 2, 0) / values.length;
const std = Math.sqrt(variance);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2) });Code explanation
- `values` is a tiny dataset that can be checked manually.
- The mean is the total divided by the number of observations.
- Variance measures average squared distance from the mean, and the square root of variance gives standard deviation.
- These summary values help you understand the scale and spread of data before choosing or evaluating a model.
Expected result: The mean and standard deviation are printed.
Practice exercise
Create a small real-world example for Distance Metrics. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.3 Euclidean Distance
Euclidean Distance (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Euclidean Distance to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Euclidean Distance
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Euclidean Distance. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.4 Manhattan Distance
Manhattan Distance (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Manhattan Distance to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Manhattan Distance
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Manhattan Distance. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.5 Choosing K
Choosing K (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Choosing K to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Choosing K
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Choosing K. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.6 KNN Classification
KNN Classification (predicting a category or class). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
An email filter receives a new message and decides whether it belongs to the 'spam' class or the 'not spam' class.
Coding example
// KNN Classification
const sigmoid = z => 1 / (1 + Math.exp(-z));
const weights = [0.8, -0.4];
const features = [2, 1];
const bias = -0.2;
const score = weights.reduce((sum, w, i) => sum + w * features[i], bias);
const probability = sigmoid(score);
const predictedClass = probability >= 0.5 ? 1 : 0;
console.log({ probability: probability.toFixed(3), predictedClass });Code explanation
- `weights`, `features`, and `bias` create a simple linear score.
- The sigmoid function converts any score into a value between 0 and 1.
- A threshold of 0.5 turns the probability into a class label.
- Printing both values helps you distinguish a model score from the final classification decision.
Expected result: A probability and a predicted class are printed.
Practice exercise
Create a second example for KNN Classification. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
21.7 KNN Regression
KNN Regression (predicting a continuous numerical value). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
A model receives the size, age, and location of a house and predicts a numerical price such as $850,000.
Coding example
// KNN Regression
const distance = (a,b) => Math.sqrt(a.reduce((s,x,i) => s + (x-b[i])**2, 0));
const items = [
{ label: 'A', x: [1, 1] },
{ label: 'B', x: [4, 4] },
{ label: 'C', x: [2, 2] }
];
const query = [2.2, 2.1];
const nearest = items.map(item => ({...item, d: distance(item.x, query)})).sort((a,b) => a.d-b.d)[0];
console.log({ nearest: nearest.label, distance: nearest.d.toFixed(3) });Code explanation
- The `distance()` function measures straight-line distance between two feature vectors.
- Each stored item has a label and a small numeric representation.
- The query is compared with every item, then results are sorted from nearest to farthest.
- The first result demonstrates how neighbor-based prediction or retrieval selects the closest example.
Expected result: The closest stored item and its distance are printed.
Practice exercise
Create a second example for KNN Regression. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
21.8 Scaling
Scaling (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Scaling to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Scaling
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Scaling. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
21.9 Weighted KNN
Weighted KNN (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
To classify a new flower, compare it with the most similar flowers already labelled. If most nearby examples are labelled 'Type A', the new flower can be classified as Type A.
Coding example
// Weighted KNN
const distance = (a,b) => Math.sqrt(a.reduce((s,x,i) => s + (x-b[i])**2, 0));
const items = [
{ label: 'A', x: [1, 1] },
{ label: 'B', x: [4, 4] },
{ label: 'C', x: [2, 2] }
];
const query = [2.2, 2.1];
const nearest = items.map(item => ({...item, d: distance(item.x, query)})).sort((a,b) => a.d-b.d)[0];
console.log({ nearest: nearest.label, distance: nearest.d.toFixed(3) });Code explanation
- The `distance()` function measures straight-line distance between two feature vectors.
- Each stored item has a label and a small numeric representation.
- The query is compared with every item, then results are sorted from nearest to farthest.
- The first result demonstrates how neighbor-based prediction or retrieval selects the closest example.
Expected result: The closest stored item and its distance are printed.
Practice exercise
Create a second example for Weighted KNN. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
21.10 Curse of Dimensionality
Curse of Dimensionality (a practical concept used within supervised learning, evaluation, and ensemble methods). Within Chapter 21, this topic connects directly to supervised learning, evaluation, and ensemble methods. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
When using this idea, separate training data from evaluation data and compare performance on examples the model did not train on. Pay attention to assumptions, important settings, error patterns, and whether the method is appropriate for classification, regression, ranking, or probability estimation.
Example
Imagine a small machine-learning project. Use Curse of Dimensionality to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Curse of Dimensionality
const rows = [[2,1],[4,2],[6,3],[8,4]];
const direction = [0.894, 0.447];
const projected = rows.map(row => row[0]*direction[0] + row[1]*direction[1]);
console.log(projected.map(x => x.toFixed(2)));Code explanation
- Each row begins with two numeric features.
- `direction` represents a chosen one-dimensional axis.
- The dot product projects each two-dimensional point onto that axis.
- The result shows how dimensionality reduction can compress several features into fewer numbers while preserving useful structure.
Expected result: One projected value is printed for each original row.
Practice exercise
Create a small real-world example for Curse of Dimensionality. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 21 Review Questions and Answers
Q1. What is Instance-Based Learning?
Answer: Instance-Based Learning is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Distance Metrics?
Answer: Distance Metrics is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Euclidean Distance?
Answer: Euclidean Distance is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Manhattan Distance?
Answer: Manhattan Distance is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Choosing K?
Answer: Choosing K is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is KNN Classification?
Answer: KNN Classification is predicting a category or class. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is KNN Regression?
Answer: KNN Regression is predicting a continuous numerical value. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Scaling?
Answer: Scaling is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Weighted KNN?
Answer: Weighted KNN is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Curse of Dimensionality?
Answer: Curse of Dimensionality is a practical concept used within supervised learning, evaluation, and ensemble methods. In this chapter, focus on the input, the method or decision, and the result that should be checked.