Chapter 12: Exploratory Data Analysis
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 15 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
12.1 Understanding a Dataset
Understanding a Dataset (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Understanding a Dataset to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Understanding a Dataset
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Understanding a Dataset. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.2 Dataset Shape
Dataset Shape (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Dataset Shape to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Dataset Shape
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Dataset Shape. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.3 Column Types
Column Types (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Column Types to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Column Types
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Column Types. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.4 Summary Statistics
Summary Statistics (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Summary Statistics to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Summary Statistics
const values = [12, 15, 11, 18, 14, 16];
const mean = values.reduce((sum, x) => sum + x, 0) / values.length;
const variance = values.reduce((sum, x) => sum + (x - mean) ** 2, 0) / values.length;
const std = Math.sqrt(variance);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2) });Code explanation
- `values` is a tiny dataset that can be checked manually.
- The mean is the total divided by the number of observations.
- Variance measures average squared distance from the mean, and the square root of variance gives standard deviation.
- These summary values help you understand the scale and spread of data before choosing or evaluating a model.
Expected result: The mean and standard deviation are printed.
Practice exercise
Create a small real-world example for Summary Statistics. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.5 Frequency Tables
Frequency Tables (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Frequency Tables to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Frequency Tables
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Frequency Tables. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.6 Histograms
Histograms (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Histograms is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Histograms
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Histograms. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
12.7 Bar Charts
Bar Charts (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Bar Charts is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Bar Charts
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Bar Charts. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
12.8 Scatter Plots
Scatter Plots (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Scatter Plots is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Scatter Plots
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Scatter Plots. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
12.9 Box Plots
Box Plots (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Box Plots is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Box Plots
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Box Plots. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
12.10 Correlation Matrices
Correlation Matrices (a standardized measure of the strength and direction of a relationship). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Correlation Matrices to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Correlation Matrices
const values = [12, 15, 11, 18, 14, 16];
const mean = values.reduce((sum, x) => sum + x, 0) / values.length;
const variance = values.reduce((sum, x) => sum + (x - mean) ** 2, 0) / values.length;
const std = Math.sqrt(variance);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2) });Code explanation
- `values` is a tiny dataset that can be checked manually.
- The mean is the total divided by the number of observations.
- Variance measures average squared distance from the mean, and the square root of variance gives standard deviation.
- These summary values help you understand the scale and spread of data before choosing or evaluating a model.
Expected result: The mean and standard deviation are printed.
Practice exercise
Create a small real-world example for Correlation Matrices. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.11 Distribution Analysis
Distribution Analysis (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Distribution Analysis to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Distribution Analysis
const outcomes = [1, 0, 1, 1, 0, 1, 0, 1];
const successes = outcomes.reduce((sum, x) => sum + x, 0);
const probability = successes / outcomes.length;
const smoothed = (successes + 1) / (outcomes.length + 2);
console.log({ probability: probability.toFixed(3), smoothed: smoothed.toFixed(3) });Code explanation
- Each `1` represents an observed success and each `0` represents a non-success.
- Dividing the number of successes by the number of observations gives an empirical probability.
- The smoothed estimate adds one pseudo-success and one pseudo-failure so very small datasets are less extreme.
- Comparing the raw and smoothed results demonstrates how probabilistic estimates can change when prior information is introduced.
Expected result: Two probability estimates are printed for comparison.
Practice exercise
Create a small real-world example for Distribution Analysis. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.12 Pattern Detection
Pattern Detection (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Pattern Detection to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Pattern Detection
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Pattern Detection. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.13 Anomaly Detection
Anomaly Detection (an observation that differs strongly from expected or normal patterns). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Anomaly Detection to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Anomaly Detection
const values = [10,11,9,12,10,11,48];
const mean = values.reduce((a,b)=>a+b,0)/values.length;
const std = Math.sqrt(values.reduce((s,x)=>s+(x-mean)**2,0)/values.length);
const flagged = values.filter(x => Math.abs((x-mean)/std) > 2);
console.log({ mean: mean.toFixed(2), std: std.toFixed(2), flagged });Code explanation
- The dataset includes one intentionally unusual value.
- Mean and standard deviation summarize the normal range of the small sample.
- Each value is converted into a standardized distance from the mean.
- Values beyond the selected threshold are flagged for investigation rather than automatically treated as errors.
Expected result: The unusual value is listed in the flagged array.
Practice exercise
Create a small real-world example for Anomaly Detection. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.14 Relationships Between Variables
Relationships Between Variables (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Relationships Between Variables to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Relationships Between Variables
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Relationships Between Variables. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
12.15 EDA Reports
EDA Reports (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 12, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use EDA Reports to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// EDA Reports
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for EDA Reports. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 12 Review Questions and Answers
Q1. What is Understanding a Dataset?
Answer: Understanding a Dataset is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Dataset Shape?
Answer: Dataset Shape is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Column Types?
Answer: Column Types is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Summary Statistics?
Answer: Summary Statistics is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Frequency Tables?
Answer: Frequency Tables is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Histograms?
Answer: Histograms is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Bar Charts?
Answer: Bar Charts is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Scatter Plots?
Answer: Scatter Plots is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Box Plots?
Answer: Box Plots is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Correlation Matrices?
Answer: Correlation Matrices is a standardized measure of the strength and direction of a relationship. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q11. What is Distribution Analysis?
Answer: Distribution Analysis is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q12. What is Pattern Detection?
Answer: Pattern Detection is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q13. What is Anomaly Detection?
Answer: Anomaly Detection is an observation that differs strongly from expected or normal patterns. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q14. What is Relationships Between Variables?
Answer: Relationships Between Variables is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q15. What is EDA Reports?
Answer: EDA Reports is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.