Redundancy
In short: Redundancy means keeping important parts of a system in duplicate so that the failure of one part does not paralyse the whole.
In more detail: A single point whose failure stops everything is called a single point of failure. Redundancy removes it by having a replacement ready.
In Depth
Examples
- Power: two power supplies, UPS and generator in the data centre.
- Network: two lines and several paths, see full mesh topology and leaf-spine.
- Storage: RAID mirrors or distributes data across several disks.
- Servers: several machines behind a load balancer. If one fails, the others take over (failover).
- Data: backups in another location. Rule: 3-2-1 (3 copies, 2 media, 1 offsite).
Active and passive
- Active-active: all work, the load is distributed.
- Active-passive: the spare waits and jumps in on failure.
Limits
Redundancy costs money and makes systems more complex. A spare that is never tested often fails when needed. Redundancy also does not replace backups: a deleted record is instantly propagated to all copies. It is a building block of IT security (availability).
See also: scaling, system architecture