Chapter 11: Data Collection
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 15 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
11.1 Structured Data
Structured Data (data organized into a consistent schema such as rows and columns). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Structured Data is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Structured Data
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Structured Data. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
11.2 Unstructured Data
Unstructured Data (data without a fixed table structure, such as text, images, audio, or video). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Unstructured Data is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Unstructured Data
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Unstructured Data. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
11.3 Tabular Data
Tabular Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small real-world project where Tabular Data is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Tabular Data
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Tabular Data. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
11.4 Text Data
Text Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Text Data to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Text Data
const documents = [
'models learn patterns from data',
'graphs connect related entities',
'retrieval finds useful context'
];
const query = 'find useful data';
const words = s => new Set(s.toLowerCase().split(/\s+/));
const q = words(query);
const scored = documents.map((text,i)=>({i,text,score:[...words(text)].filter(w=>q.has(w)).length})).sort((a,b)=>b.score-a.score);
console.log(scored[0]);Code explanation
- The documents and query are converted into simple sets of lowercase words.
- Each document receives one point for every word it shares with the query.
- Sorting by score produces a basic relevance ranking.
- Real retrieval systems use stronger representations and indexes, but this tiny example makes the retrieval step visible.
Expected result: The highest-scoring document is printed.
Practice exercise
Create a small real-world example for Text Data. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.5 Image Data
Image Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Image Data to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Image Data
const signal = [1,2,3,4,3,2,1];
const kernel = [1,0,-1];
const result = [];
for(let i=0;i<=signal.length-kernel.length;i++){
result.push(kernel.reduce((s,k,j)=>s+k*signal[i+j],0));
}
console.log(result);Code explanation
- `signal` is a tiny stand-in for a row of pixel or sensor values.
- `kernel` is a small filter that is moved across the signal.
- At each position, neighboring values are multiplied by kernel weights and summed.
- The output highlights local changes, illustrating the main operation behind convolutional feature extraction.
Expected result: A short filtered feature sequence is printed.
Practice exercise
Create a small real-world example for Image Data. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.6 Audio Data
Audio Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Audio Data to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Audio Data
const sourceA = [0.7, 0.2, 0.5];
const sourceB = [0.1, 0.9];
const combined = [...sourceA, ...sourceB];
const score = combined.reduce((s,x)=>s+x,0) / combined.length;
console.log({ combined, score: score.toFixed(3) });Code explanation
- Two different feature sources are represented by separate numeric vectors.
- The spread operator combines them into one representation.
- A simple average produces one downstream score from the fused features.
- Real multimodal or scientific systems usually learn how much weight each source deserves, but the example shows the fusion step clearly.
Expected result: The combined feature vector and a summary score are printed.
Practice exercise
Create a small real-world example for Audio Data. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.7 Video Data
Video Data (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Video Data to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Video Data
const signal = [1,2,3,4,3,2,1];
const kernel = [1,0,-1];
const result = [];
for(let i=0;i<=signal.length-kernel.length;i++){
result.push(kernel.reduce((s,k,j)=>s+k*signal[i+j],0));
}
console.log(result);Code explanation
- `signal` is a tiny stand-in for a row of pixel or sensor values.
- `kernel` is a small filter that is moved across the signal.
- At each position, neighboring values are multiplied by kernel weights and summed.
- The output highlights local changes, illustrating the main operation behind convolutional feature extraction.
Expected result: A short filtered feature sequence is printed.
Practice exercise
Create a small real-world example for Video Data. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.8 Time-Series Data
Time-Series Data (data recorded in time order). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
A store records daily sales for two years. A forecasting model uses trend, seasonality, and recent sales to estimate sales for the next week.
Coding example
// Time-Series Data
const sequence = [2,4,3,5,7];
let state = 0;
const alpha = 0.6;
const states = sequence.map(x => {
state = alpha * x + (1-alpha) * state;
return Number(state.toFixed(2));
});
console.log(states);Code explanation
- The input values arrive in order, so earlier information can influence later calculations.
- `state` stores a running memory instead of treating every value independently.
- The update blends the new input with the previous state.
- The printed states demonstrate the idea of sequential models and online updates maintaining information through time.
Expected result: A state value is printed for every step in the sequence.
Practice exercise
Create a second example for Time-Series Data. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
11.9 APIs
APIs (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use APIs to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// APIs
const model = input => input.reduce((a,b)=>a+b,0) / input.length;
const cache = new Map();
function predict(input){
const key = JSON.stringify(input);
if(cache.has(key)) return { value: cache.get(key), cached: true };
const value = model(input); cache.set(key,value);
return { value, cached: false };
}
console.log(predict([2,4,6]));
console.log(predict([2,4,6]));Code explanation
- `model()` stands in for a trained prediction function.
- `predict()` creates a stable key from the request so repeated inputs can be recognized.
- The first request computes and stores the result; the second request reuses it.
- This demonstrates a production concern—serving predictions efficiently—without depending on any particular deployment vendor.
Expected result: The first result is uncached and the second is returned from the cache.
Practice exercise
Create a small real-world example for APIs. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.10 Databases
Databases (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Databases to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Databases
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Databases. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.11 Sensors
Sensors (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Sensors to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Sensors
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Sensors. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.12 Public Datasets
Public Datasets (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Public Datasets to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Public Datasets
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Public Datasets. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.13 Sampling
Sampling (selecting a subset of a larger population or dataset). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Sampling to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Sampling
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Sampling. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.14 Data Quality
Data Quality (a practical concept used within data collection, understanding, cleaning, and preparation). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Data Quality to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Data Quality
const rows = [
{ age: 22, score: 71 },
{ age: null, score: 88 },
{ age: 35, score: 93 }
];
const knownAges = rows.filter(r => r.age !== null).map(r => r.age);
const fallbackAge = knownAges.reduce((a,b) => a+b, 0) / knownAges.length;
const cleaned = rows.map(r => ({ ...r, age: r.age ?? fallbackAge }));
console.log(cleaned);Code explanation
- The sample rows deliberately contain one missing value so you can see a preprocessing decision.
- Known ages are separated and averaged to create a simple fallback value.
- `map()` builds a new cleaned dataset instead of modifying the original rows in place.
- The final log lets you verify that every row now has a usable numeric age.
Expected result: A cleaned array is printed with the missing age filled.
Practice exercise
Create a small real-world example for Data Quality. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
11.15 Data Provenance
Data Provenance (information about where data came from and how it changed). Within Chapter 11, this topic connects directly to data collection, understanding, cleaning, and preparation. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
In a real project, this step can affect every model that comes later. Check data types, missing values, scale, categories, unusual records, and whether the transformation can be repeated consistently on new data. Good preparation reduces avoidable errors and helps make evaluation more trustworthy.
Example
Imagine a small machine-learning project. Use Data Provenance to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Data Provenance
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Data Provenance. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
Chapter 11 Review Questions and Answers
Q1. What is Structured Data?
Answer: Structured Data is data organized into a consistent schema such as rows and columns. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Unstructured Data?
Answer: Unstructured Data is data without a fixed table structure, such as text, images, audio, or video. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Tabular Data?
Answer: Tabular Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Text Data?
Answer: Text Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Image Data?
Answer: Image Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Audio Data?
Answer: Audio Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Video Data?
Answer: Video Data is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Time-Series Data?
Answer: Time-Series Data is data recorded in time order. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is APIs?
Answer: APIs is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Databases?
Answer: Databases is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q11. What is Sensors?
Answer: Sensors is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q12. What is Public Datasets?
Answer: Public Datasets is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q13. What is Sampling?
Answer: Sampling is selecting a subset of a larger population or dataset. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q14. What is Data Quality?
Answer: Data Quality is a practical concept used within data collection, understanding, cleaning, and preparation. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q15. What is Data Provenance?
Answer: Data Provenance is information about where data came from and how it changed. In this chapter, focus on the input, the method or decision, and the result that should be checked.