Chapter 62: High Availability, Disaster Recovery, and Business Continuity
This chapter follows the topics shown in the Networking chapter menu. Work through each section in order, then use the review questions to check recall and troubleshooting reasoning.
62.1 Single Point of Failure
A single point of failure is a component whose failure can interrupt the entire service or path because no effective alternative exists.
Scenario: a user reports a problem related to Single Point of Failure. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.2 Redundancy
Redundancy is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.3 High Availability
High availability reduces service interruption through redundancy, failure detection, failover, resilient design, and regular verification.
Scenario: a user reports a problem related to High Availability. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.4 Device Redundancy
Device Redundancy is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Device Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.5 Link Redundancy
Link Redundancy is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Link Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.6 ISP Redundancy
ISP Redundancy is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to ISP Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.7 Power Redundancy
Power Redundancy is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Power Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.8 UPS
UPS is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to UPS. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.9 Generators
Generators is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Generators. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.10 Blackouts
Blackouts is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Blackouts. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.11 Brownouts
Brownouts is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Brownouts. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.12 Surges
Surges is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Surges. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.13 Spikes
Spikes is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Spikes. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.14 Disaster Recovery
Disaster Recovery is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Disaster Recovery. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.15 Business Continuity
Business Continuity is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Business Continuity. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.16 Full Backups
Full Backups is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Full Backups. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.17 Incremental Backups
Incremental Backups is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Incremental Backups. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.18 Differential Backups
Differential Backups is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Differential Backups. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.19 RTO
RTO is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to RTO. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.20 RPO
RPO is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to RPO. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.21 MTTR
MTTR is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to MTTR. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.22 MTBF
MTBF is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to MTBF. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.23 Hot Sites
Hot Sites is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Hot Sites. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.24 Warm Sites
Warm Sites is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Warm Sites. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.25 Cold Sites
Cold Sites is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Cold Sites. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.26 Active/Active
Active/Active is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Active/Active. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.27 Active/Passive
Active/Passive is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Active/Passive. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.28 Failover
Failover is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Failover. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.29 Failback
Failback is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Failback. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.30 FHRP
FHRP is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to FHRP. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.31 HSRP
HSRP is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to HSRP. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.32 VRRP
VRRP is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to VRRP. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.33 Backup Testing
Backup Testing is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Backup Testing. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
62.34 DR Testing
DR Testing is one of the core topics in High Availability, Disaster Recovery, and Business Continuity. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to DR Testing. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.