EASYTUTORGUIDE

Practical tutorials, tools, courses, digital skills, and business promotion.

Free Learning
Google Translate

Chapter 16: Data Preprocessing

Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.

Beginner FriendlyExamplesPracticeExpert Topics
Estimated reading time0% read

What this chapter covers

This chapter contains 15 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.

16.1 Train/Test Splitting

Train/Test Splitting (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Train/Test Splitting is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Train/Test Splitting
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a second example for Train/Test Splitting. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

16.2 Validation Sets

Validation Sets (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Validation Sets to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Validation Sets
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Validation Sets. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.3 Scaling

Scaling (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Scaling is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Scaling
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a second example for Scaling. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

16.4 Standardization

Standardization (rescaling values using the mean and standard deviation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Standardization is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Standardization
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a second example for Standardization. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

16.5 Normalization

Normalization (changing values to a common numerical scale). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Normalization is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Normalization
const a = [2, 4, 6];
const b = [1, 3, 5];

const dot = a.reduce((sum, value, i) => sum + value * b[i], 0);
const magnitude = Math.sqrt(a.reduce((sum, value) => sum + value ** 2, 0));

console.log({ dot, magnitude: magnitude.toFixed(2) });

Code explanation

  1. The arrays `a` and `b` represent small numeric vectors so the calculation stays easy to inspect.
  2. `reduce()` walks through the values and combines them into one result, which is useful for many linear-algebra operations.
  3. The magnitude calculation squares each value, adds the squares, and takes the square root.
  4. The final object prints values you can compare by hand before using the same idea with larger data.

Expected result: A dot-product value and a vector magnitude are printed.

Practice exercise

Create a second example for Normalization. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

16.6 Encoding Categories

Encoding Categories (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Encoding Categories to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Encoding Categories
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Encoding Categories. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.7 One-Hot Encoding

One-Hot Encoding (representing categories with separate zero-or-one indicator columns). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where One-Hot Encoding is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// One-Hot Encoding
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a second example for One-Hot Encoding. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

16.8 Ordinal Encoding

Ordinal Encoding (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Ordinal Encoding to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Ordinal Encoding
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Ordinal Encoding. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.9 Label Encoding

Label Encoding (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Label Encoding to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Label Encoding
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Label Encoding. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.10 Log Transformations

Log Transformations (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Log Transformations to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Log Transformations
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Log Transformations. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.11 Power Transformations

Power Transformations (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Power Transformations to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Power Transformations
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Power Transformations. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.12 Imbalanced Data

Imbalanced Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Imbalanced Data to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Imbalanced Data
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Imbalanced Data. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.13 Resampling

Resampling (selecting a subset of a larger population or dataset). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Resampling to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Resampling
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Resampling. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.14 Data Pipelines

Data Pipelines (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Data Pipelines to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Data Pipelines
const tools = {
  average: values => values.reduce((a,b)=>a+b,0)/values.length,
  maximum: values => Math.max(...values)
};
const task = { tool: 'average', input: [4,7,9,10] };
const result = tools[task.tool](task.input);

console.log({ task, result });

Code explanation

  1. The `tools` object acts as a small registry of allowed operations.
  2. The task explicitly names which tool should run and provides its input.
  3. The dispatcher selects the requested function and executes it.
  4. This pattern demonstrates controlled tool use and workflow orchestration without giving unrestricted access to arbitrary operations.

Expected result: The selected tool and its computed result are printed.

Practice exercise

Create a small real-world example for Data Pipelines. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

16.15 Preventing Leakage

Preventing Leakage (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 16, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Preventing Leakage to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Preventing Leakage
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Preventing Leakage. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

Chapter 16 Review Questions and Answers

Q1. What is Train/Test Splitting?

Answer: Train/Test Splitting is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q2. What is Validation Sets?

Answer: Validation Sets is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q3. What is Scaling?

Answer: Scaling is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q4. What is Standardization?

Answer: Standardization is rescaling values using the mean and standard deviation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q5. What is Normalization?

Answer: Normalization is changing values to a common numerical scale. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q6. What is Encoding Categories?

Answer: Encoding Categories is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q7. What is One-Hot Encoding?

Answer: One-Hot Encoding is representing categories with separate zero-or-one indicator columns. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q8. What is Ordinal Encoding?

Answer: Ordinal Encoding is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q9. What is Label Encoding?

Answer: Label Encoding is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q10. What is Log Transformations?

Answer: Log Transformations is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q11. What is Power Transformations?

Answer: Power Transformations is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q12. What is Imbalanced Data?

Answer: Imbalanced Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q13. What is Resampling?

Answer: Resampling is selecting a subset of a larger population or dataset. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q14. What is Data Pipelines?

Answer: Data Pipelines is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q15. What is Preventing Leakage?

Answer: Preventing Leakage is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.