EASYTUTORGUIDE

Practical tutorials, tools, courses, digital skills, and business promotion.

Free Learning
Google Translate

Chapter 13: Data Cleaning

Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.

Beginner FriendlyExamplesPracticeExpert Topics
Estimated reading time0% read

What this chapter covers

This chapter contains 15 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.

13.1 Missing Data

Missing Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Missing Data is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Missing Data
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a second example for Missing Data. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

13.2 Detecting Missing Values

Detecting Missing Values (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Detecting Missing Values is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Detecting Missing Values
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a second example for Detecting Missing Values. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

13.3 Dropping Missing Values

Dropping Missing Values (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Dropping Missing Values is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Dropping Missing Values
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a second example for Dropping Missing Values. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

13.4 Imputation

Imputation (filling in missing data with a chosen value or estimate). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Imputation is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Imputation
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a second example for Imputation. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

13.5 Duplicate Records

Duplicate Records (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small real-world project where Duplicate Records is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.

Coding example

// Duplicate Records
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a second example for Duplicate Records. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.

13.6 Incorrect Data Types

Incorrect Data Types (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Incorrect Data Types to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Incorrect Data Types
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Incorrect Data Types. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.7 Invalid Values

Invalid Values (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Invalid Values to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Invalid Values
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Invalid Values. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.8 Outliers

Outliers (an observation that is unusually different from most other observations). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Outliers to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Outliers
const values = [10,11,9,12,10,11,48];
const mean = values.reduce((a,b)=>a+b,0)/values.length;
const std = Math.sqrt(values.reduce((s,x)=>s+(x-mean)**2,0)/values.length);
const flagged = values.filter(x => Math.abs((x-mean)/std) > 2);

console.log({ mean: mean.toFixed(2), std: std.toFixed(2), flagged });

Code explanation

  1. The dataset includes one intentionally unusual value.
  2. Mean and standard deviation summarize the normal range of the small sample.
  3. Each value is converted into a standardized distance from the mean.
  4. Values beyond the selected threshold are flagged for investigation rather than automatically treated as errors.

Expected result: The unusual value is listed in the flagged array.

Practice exercise

Create a small real-world example for Outliers. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.9 Inconsistent Categories

Inconsistent Categories (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Inconsistent Categories to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Inconsistent Categories
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Inconsistent Categories. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.10 Text Cleaning

Text Cleaning (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Text Cleaning to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Text Cleaning
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a small real-world example for Text Cleaning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.11 Date Cleaning

Date Cleaning (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Date Cleaning to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Date Cleaning
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a small real-world example for Date Cleaning. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.12 Data Validation

Data Validation (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Data Validation to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Data Validation
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a small real-world example for Data Validation. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.13 Data Leakage

Data Leakage (using information during training that would not truly be available at prediction time). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Data Leakage to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Data Leakage
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);

console.log(results);

Code explanation

  1. The sample starts with a small list of inputs so every result can be checked manually.
  2. `transform()` represents the main operation for this topic in a deliberately simple form.
  3. `map()` applies the same rule consistently to every item and returns a new result array.
  4. Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.

Expected result: A transformed result is printed for each input value.

Practice exercise

Create a small real-world example for Data Leakage. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.14 Data Quality Checks

Data Quality Checks (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Data Quality Checks to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Data Quality Checks
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a small real-world example for Data Quality Checks. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

13.15 Cleaning Pipelines

Cleaning Pipelines (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 13, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.

In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.

Example

Imagine a small machine-learning project. Use Cleaning Pipelines to decide what information is needed, what step happens next, and what result should be checked.

Coding example

// Cleaning Pipelines
const rows = [
  { age: 22, score: 71 },
  { age: null, score: 88 },
  { age: 35, score: 93 }
];

const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));

console.log(cleaned);

Code explanation

  1. The sample rows deliberately contain one missing value so you can see a preprocessing decision.
  2. Known ages are separated and averaged to create a simple fallback value.
  3. `map()` builds a new cleaned dataset instead of modifying the original rows in place.
  4. The final log lets you verify that every row now has a usable numeric age.

Expected result: A cleaned array is printed with the missing age filled.

Practice exercise

Create a small real-world example for Cleaning Pipelines. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.

Chapter 13 Review Questions and Answers

Q1. What is Missing Data?

Answer: Missing Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q2. What is Detecting Missing Values?

Answer: Detecting Missing Values is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q3. What is Dropping Missing Values?

Answer: Dropping Missing Values is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q4. What is Imputation?

Answer: Imputation is filling in missing data with a chosen value or estimate. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q5. What is Duplicate Records?

Answer: Duplicate Records is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q6. What is Incorrect Data Types?

Answer: Incorrect Data Types is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q7. What is Invalid Values?

Answer: Invalid Values is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q8. What is Outliers?

Answer: Outliers is an observation that is unusually different from most other observations. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q9. What is Inconsistent Categories?

Answer: Inconsistent Categories is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q10. What is Text Cleaning?

Answer: Text Cleaning is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q11. What is Date Cleaning?

Answer: Date Cleaning is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q12. What is Data Validation?

Answer: Data Validation is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q13. What is Data Leakage?

Answer: Data Leakage is using information during training that would not truly be available at prediction time. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q14. What is Data Quality Checks?

Answer: Data Quality Checks is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.

Q15. What is Cleaning Pipelines?

Answer: Cleaning Pipelines is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.