High availability (HA) means designing the network so that a single failure, such as a dead link, a failed supervisor or a crashed gateway, causes little or no disruption. You get there with redundancy at several levels: redundant links, redundant devices, redundant components inside a device, and protocols that fail over quickly between them.
Link and device redundancy is the foundation. Access switches uplink to two distribution switches, distribution switches connect to two core switches, and critical servers are dual-homed. EtherChannel bundles several physical links into one logical link so losing one member does not change the topology. Redundant paths only help if something decides quickly which path to use, which is the job of spanning tree at Layer 2 and routing protocols at Layer 3.
Endpoints are the weak spot. A PC is configured with a single default gateway IP address and has no routing protocol to discover an alternative. First hop redundancy protocols (FHRPs) solve this by letting two or more routers share a virtual IP address and a virtual MAC address. Hosts use the virtual IP as their gateway. One router actively forwards traffic for it; if that router fails, another takes over the same virtual addresses and the hosts never notice. Hot Standby Router Protocol (HSRP) is Cisco proprietary and uses one active and one standby router. Virtual Router Redundancy Protocol (VRRP) is an open standard with one master and one or more backups. Gateway Load Balancing Protocol (GLBP) is Cisco proprietary and lets several routers forward at once by handing out different virtual MAC addresses to different hosts. Features such as priority, preemption and interface or object tracking let you control which router is active and move the role if an uplink fails.
Inside a chassis switch or router with two route processors or supervisors, you also need redundancy. Stateful switchover (SSO) keeps the standby supervisor synchronized with the active one: configuration, and state information such as interface and Layer 2 protocol state. If the active supervisor fails, the standby takes over without resetting line cards, so forwarding in hardware continues. SSO on its own does not preserve routing protocol adjacencies; the new supervisor must rebuild them. That is why SSO is paired with Nonstop Forwarding (NSF), which lets the device keep forwarding using the existing forwarding table while graceful restart helpers (the neighbors) keep their adjacencies and routes in place during the rebuild.
Stacking and virtual switching add another layer. Catalyst switch stacks and StackWise Virtual let two or more physical switches act as one logical switch, with one control plane and SSO between members. Neighbors can then use a multichassis EtherChannel to both physical switches, removing spanning tree blocked links.
When you design HA, remember that more redundancy adds complexity. Too many parallel paths can slow convergence and make troubleshooting harder. The usual guidance is two of everything at each layer, with fast, deterministic failover.
Key terms
- FHRP
- First hop redundancy protocol: lets several routers share a virtual gateway IP and MAC so hosts survive a gateway failure.
- HSRP
- Hot Standby Router Protocol, a Cisco FHRP with one active and one standby router per group.
- VRRP
- Virtual Router Redundancy Protocol, an open-standard FHRP with a master and backups.
- SSO
- Stateful switchover: the standby supervisor stays synchronized and takes over without resetting line cards.
- NSF
- Nonstop Forwarding: keeps forwarding packets using existing forwarding tables while routing protocols reconverge after a switchover.
A distribution pair runs HSRP for each user VLAN, with switch A at priority 110 and preemption enabled and tracking its core uplink. When A's uplink fails, tracking lowers its priority below B's, B becomes active, and users keep using the same gateway address without noticing.
Check yourself
Why do hosts need an FHRP when the network already has two gateway routers?
Hosts have a single statically configured or DHCP-assigned gateway and cannot detect its failure; an FHRP presents one virtual IP and MAC that another router takes over.
Which FHRP lets several routers forward traffic for the same group at the same time?
GLBP, by assigning different virtual MAC addresses to different hosts.
What does SSO synchronize to the standby supervisor?
Configuration and state information, so the standby can take over without resetting line cards; NSF is added so forwarding continues while routing protocols rebuild.