Chapter 50: WAN, VPN, Cloud, and HA Troubleshooting
This chapter follows the topics shown in the Networking chapter menu. Work through each section in order, then use the review questions to check recall and troubleshooting reasoning.
50.1 WAN Failure Symptoms
WAN Failure Symptoms is a troubleshooting condition. The useful approach is to confirm symptoms, determine scope, identify the relevant layer, compare actual values with the intended design, test one theory at a time, and verify service after the fix.
Scenario: a user reports a problem related to WAN Failure Symptoms. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.2 Demarc Troubleshooting
Demarc Troubleshooting is a troubleshooting condition. The useful approach is to confirm symptoms, determine scope, identify the relevant layer, compare actual values with the intended design, test one theory at a time, and verify service after the fix.
Scenario: a user reports a problem related to Demarc Troubleshooting. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.3 CPE Troubleshooting
CPE Troubleshooting is a troubleshooting condition. The useful approach is to confirm symptoms, determine scope, identify the relevant layer, compare actual values with the intended design, test one theory at a time, and verify service after the fix.
Scenario: a user reports a problem related to CPE Troubleshooting. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.4 ISP Outages
ISP Outages is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to ISP Outages. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.5 WAN Congestion
WAN Congestion is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to WAN Congestion. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.6 Redundant WANs
Redundant WANs is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Redundant WANs. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.7 Failover
Failover is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Failover. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.8 Failback
Failback is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Failback. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.9 Failover Failure
Failover Failure is a troubleshooting condition. The useful approach is to confirm symptoms, determine scope, identify the relevant layer, compare actual values with the intended design, test one theory at a time, and verify service after the fix.
Scenario: a user reports a problem related to Failover Failure. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.10 Remote-Access VPN
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.11 Site-to-Site VPN
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.12 VPN Authentication
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.13 Tunnel-Up/No-Traffic Problems
Tunnel-Up/No-Traffic Problems is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Tunnel-Up/No-Traffic Problems. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.14 Split Tunneling
Split Tunneling is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Split Tunneling. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.15 Full Tunneling
Full Tunneling is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Full Tunneling. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.16 VPN DNS
DNS translates names into resource records such as IP addresses, aliases, mail-routing information, and service data. Client caching and TTL values affect how quickly changes become visible.
Example: a user can reach 203.0.113.20 but cannot reach server.example by name. That difference points toward name resolution, DNS reachability, record content, cache state, or search-suffix behavior rather than basic IP routing.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.17 VPN MTU
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.18 VPN Latency
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.19 VPN Concentrator Capacity
A VPN creates a protected logical connection across an untrusted or shared network. The design must define endpoints, authentication, encryption, routing, and failure behavior.
Example: two sites can reach the public internet but cannot pass private traffic through the secure tunnel. Check peer reachability, negotiation state, authentication, encryption proposals, interesting traffic, NAT interaction, routes, and policy.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.20 SD-WAN Underlay
SD-WAN uses centralized policy and an overlay to steer application traffic across multiple WAN transports according to availability, latency, loss, jitter, and business intent.
Scenario: a user reports a problem related to SD-WAN Underlay. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.21 SD-WAN Overlay
SD-WAN uses centralized policy and an overlay to steer application traffic across multiple WAN transports according to availability, latency, loss, jitter, and business intent.
Scenario: a user reports a problem related to SD-WAN Overlay. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.22 SD-WAN Path Selection
PAT lets many inside hosts share one or a small number of public IPv4 addresses by tracking transport protocol and port mappings.
Scenario: a user reports a problem related to SD-WAN Path Selection. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.23 MPLS
MPLS is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to MPLS. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.24 Cloud Connectivity
Cloud Connectivity is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Cloud Connectivity. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.25 Cloud Virtual Networks
Cloud Virtual Networks is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Cloud Virtual Networks. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.26 Cloud Route Tables
Cloud Route Tables affects how Layer 3 devices choose a path toward destination networks. Correct routing depends on the destination prefix, route source, next hop or exit interface, route preference, metric, and reachability of the next step.
Example: a router receives a packet for 10.20.30.40 and has several matching routes. It selects the most specific matching prefix, then forwards toward the route's next hop or exit interface if that path is usable.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.27 Cloud Security Rules
Cloud Security Rules is a network-security concept. A sound design identifies the protected asset, trust boundary, possible abuse path, preventive controls, detection signals, and recovery steps.
Scenario: a user reports a problem related to Cloud Security Rules. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.28 Cloud NAT
NAT changes IP address information as traffic crosses a translation boundary. PAT is a many-to-one form that also distinguishes conversations by transport-layer port numbers.
Scenario: a user reports a problem related to Cloud NAT. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.29 Hybrid Cloud
Hybrid Cloud is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Hybrid Cloud. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.30 Overlapping Addresses
Ping is a basic reachability and round-trip-time test. A failed ping does not always prove the destination is down because policy may block ICMP.
Example: ping the local loopback, local interface, default gateway, remote IP, and finally a hostname. The first failed step helps narrow the fault domain, but remember that ICMP filtering can produce false negatives.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
Command or data example
ping 192.0.2.1
50.31 Hybrid DNS
DNS translates names into resource records such as IP addresses, aliases, mail-routing information, and service data. Client caching and TTL values affect how quickly changes become visible.
Example: a user can reach 203.0.113.20 but cannot reach server.example by name. That difference points toward name resolution, DNS reachability, record content, cache state, or search-suffix behavior rather than basic IP routing.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.32 High Availability
High availability reduces service interruption through redundancy, failure detection, failover, resilient design, and regular verification.
Scenario: a user reports a problem related to High Availability. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.33 Redundancy
Redundancy is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Redundancy. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.34 Single Points of Failure
Single Points of Failure is a troubleshooting condition. The useful approach is to confirm symptoms, determine scope, identify the relevant layer, compare actual values with the intended design, test one theory at a time, and verify service after the fix.
Scenario: a user reports a problem related to Single Points of Failure. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.35 Active/Passive
Active/Passive is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Active/Passive. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.36 Active/Active
Active/Active is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Active/Active. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.37 Heartbeats
Heartbeats is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Heartbeats. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.38 Split-Brain
Split-Brain is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to Split-Brain. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.39 UPS/Generator
UPS/Generator is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to UPS/Generator. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.
50.40 RTO/RPO
RTO/RPO is one of the core topics in WAN, VPN, Cloud, and HA Troubleshooting. Understand what the term represents, where it operates in the network, what information it uses, and what observable behavior confirms that it is working correctly.
Scenario: a user reports a problem related to RTO/RPO. Start by writing the exact symptom, affected users, first known failure time, and recent changes. Then choose the smallest safe test that can prove or disprove one cause.
What to check
- Confirm the symptom and determine whether the problem affects one host, one segment, one site, or many sites.
- Compare actual configuration and measurements with the intended design, baseline, or documentation.
- Change one variable at a time, verify the result, and document both the cause and the final fix.