Chapter 54: Reinforcement Learning
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 14 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
54.1 Agents
Agents (a system that observes, decides, and acts toward a goal). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
An AI agent receives a goal, decides which steps are needed, uses an available tool when necessary, checks the result, and continues until the task is complete.
Coding example
// Agents
const tools = {
average: values => values.reduce((a,b)=>a+b,0)/values.length,
maximum: values => Math.max(...values)
};
const task = { tool: 'average', input: [4,7,9,10] };
const result = tools[task.tool](task.input);
console.log({ task, result });Code explanation
- The `tools` object acts as a small registry of allowed operations.
- The task explicitly names which tool should run and provides its input.
- The dispatcher selects the requested function and executes it.
- This pattern demonstrates controlled tool use and workflow orchestration without giving unrestricted access to arbitrary operations.
Expected result: The selected tool and its computed result are printed.
Practice exercise
Create a second example for Agents. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.2 Environments
Environments (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small real-world project where Environments is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Environments
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Environments. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.3 States
States (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small real-world project where States is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// States
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for States. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.4 Actions
Actions (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small real-world project where Actions is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Actions
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Actions. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.5 Rewards
Rewards (feedback indicating how desirable an agent outcome is). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
An agent moves through a small maze. Reaching the goal gives a positive reward, hitting a blocked path gives a penalty, and repeated attempts help the agent learn better actions.
Coding example
// Rewards
let qValue = 0.4;
const reward = 1;
const nextBest = 0.7;
const rate = 0.2;
const discount = 0.9;
qValue = qValue + rate * (reward + discount * nextBest - qValue);
console.log(qValue.toFixed(3));Code explanation
- `qValue` is the current estimate of how useful an action is.
- The reward represents immediate feedback from the environment.
- The next-state estimate is discounted because future rewards are usually treated as less certain.
- The update moves the old estimate partway toward the new target instead of replacing it all at once.
Expected result: The updated action-value estimate is printed.
Practice exercise
Create a second example for Rewards. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.6 Policies
Policies (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small machine-learning project. Use Policies to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Policies
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Policies. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
54.7 Episodes
Episodes (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small real-world project where Episodes is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Episodes
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Episodes. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.8 Markov Decision Processes
Markov Decision Processes (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small machine-learning project. Use Markov Decision Processes to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Markov Decision Processes
const outcomes = [1, 0, 1, 1, 0, 1, 0, 1];
const successes = outcomes.reduce((sum, x) => sum + x, 0);
const probability = successes / outcomes.length;
const smoothed = (successes + 1) / (outcomes.length + 2);
console.log({ probability: probability.toFixed(3), smoothed: smoothed.toFixed(3) });Code explanation
- Each `1` represents an observed success and each `0` represents a non-success.
- Dividing the number of successes by the number of observations gives an empirical probability.
- The smoothed estimate adds one pseudo-success and one pseudo-failure so very small datasets are less extreme.
- Comparing the raw and smoothed results demonstrates how probabilistic estimates can change when prior information is introduced.
Expected result: Two probability estimates are printed for comparison.
Practice exercise
Create a small real-world example for Markov Decision Processes. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
54.9 Value Functions
Value Functions (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small machine-learning project. Use Value Functions to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Value Functions
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Value Functions. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
54.10 Q-Learning
Q-Learning (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
An agent moves through a small maze. Reaching the goal gives a positive reward, hitting a blocked path gives a penalty, and repeated attempts help the agent learn better actions.
Coding example
// Q-Learning
let qValue = 0.4;
const reward = 1;
const nextBest = 0.7;
const rate = 0.2;
const discount = 0.9;
qValue = qValue + rate * (reward + discount * nextBest - qValue);
console.log(qValue.toFixed(3));Code explanation
- `qValue` is the current estimate of how useful an action is.
- The reward represents immediate feedback from the environment.
- The next-state estimate is discounted because future rewards are usually treated as less certain.
- The update moves the old estimate partway toward the new target instead of replacing it all at once.
Expected result: The updated action-value estimate is printed.
Practice exercise
Create a second example for Q-Learning. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.11 Deep Q Networks
Deep Q Networks (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small machine-learning project. Use Deep Q Networks to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Deep Q Networks
const graph = { A:['B','C'], B:['D'], C:['D'], D:[] };
const visited = new Set();
const queue = ['A'];
while(queue.length){
const node = queue.shift();
if(visited.has(node)) continue;
visited.add(node);
queue.push(...graph[node]);
}
console.log([...visited]);Code explanation
- The object stores a small graph as a list of neighbors for each node.
- A queue starts from node A and explores connected nodes breadth-first.
- The `visited` set prevents repeated work when different paths reach the same node.
- This traversal pattern is a foundation for graph features, connectivity checks, and many graph-learning workflows.
Expected result: The reachable nodes are printed in traversal order.
Practice exercise
Create a small real-world example for Deep Q Networks. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
54.12 Policy Gradients
Policy Gradients (a vector showing the direction and rate of fastest increase of a function). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
An agent moves through a small maze. Reaching the goal gives a positive reward, hitting a blocked path gives a penalty, and repeated attempts help the agent learn better actions.
Coding example
// Policy Gradients
const loss = x => (x - 6) ** 2;
const derivative = x => (loss(x + 0.0001) - loss(x - 0.0001)) / 0.0002;
let value = 0;
const rate = 0.1;
for (let step = 0; step < 6; step++) {
value -= rate * derivative(value);
}
console.log({ value: value.toFixed(3), loss: loss(value).toFixed(3) });Code explanation
- `loss()` gives a simple objective: values closer to the target produce a smaller error.
- `derivative()` estimates the slope by checking the loss just to the left and right of the current value.
- The loop repeatedly moves the value opposite the slope, which demonstrates the core idea behind gradient-based optimization.
- Printing both the final value and loss lets you confirm that the search moved toward a better solution.
Expected result: The value moves toward the target and the loss becomes smaller.
Practice exercise
Create a second example for Policy Gradients. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
54.13 Actor-Critic Methods
Actor-Critic Methods (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small machine-learning project. Use Actor-Critic Methods to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Actor-Critic Methods
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Actor-Critic Methods. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
54.14 Exploration vs Exploitation
Exploration vs Exploitation (a practical concept used within reinforcement learning). Within Chapter 54, this topic connects directly to reinforcement learning. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
The input may be images, sequences, graphs, interactions, rewards, or multiple modalities. Define the task and evaluation measure first, then check whether the representation and model architecture preserve the information needed for the final decision or generated output.
Example
Imagine a small real-world project where Exploration vs Exploitation is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Exploration vs Exploitation
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a second example for Exploration vs Exploitation. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
Chapter 54 Review Questions and Answers
Q1. What is Agents?
Answer: Agents is a system that observes, decides, and acts toward a goal. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Environments?
Answer: Environments is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is States?
Answer: States is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Actions?
Answer: Actions is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Rewards?
Answer: Rewards is feedback indicating how desirable an agent outcome is. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Policies?
Answer: Policies is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Episodes?
Answer: Episodes is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Markov Decision Processes?
Answer: Markov Decision Processes is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Value Functions?
Answer: Value Functions is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q10. What is Q-Learning?
Answer: Q-Learning is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q11. What is Deep Q Networks?
Answer: Deep Q Networks is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q12. What is Policy Gradients?
Answer: Policy Gradients is a vector showing the direction and rate of fastest increase of a function. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q13. What is Actor-Critic Methods?
Answer: Actor-Critic Methods is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q14. What is Exploration vs Exploitation?
Answer: Exploration vs Exploitation is a practical concept used within reinforcement learning. In this chapter, focus on the input, the method or decision, and the result that should be checked.