Glossary · System coordination, integration and orchestration
Failover
Also known as: switchover, automatic failover
German: Failover
In high-availability systems, failover is the automatic or manual transfer of a function, service or workload from a failed or unavailable component to a redundant standby component, so that the function continues with minimal interruption.
- System integration
In one sentence
Failover transfers a function from a failed component to a redundant standby so that operation continues with minimal interruption.
Example
When the primary SCADA server stops responding, the standby server takes over the client connections and the operators continue working after a short interruption.
How it applies
- Engineering: Failover needs failure detection, typically via Heartbeat monitoring, a decision rule and state transfer. Define what may be lost during failover (for example data in transit) and how clients reconnect.
- Risks: A wrong failure detection can cause “split brain,” where both units believe they are active. Designs use quorum, fencing or leases (see Lease (distributed systems)) to prevent this.
- Maintenance: Failover paths are used rarely and can fail unnoticed. Test them periodically and after changes.
- Documentation: Operating and service documentation should describe how failover is indicated, what operators must check afterwards and how to switch back.
Failover vs. failback
Failover moves the function to the standby. Failback moves it back to the repaired primary; it is usually a planned, manual step to avoid switching back and forth.