Feature Flags (Kill Switch)
In short: A switch that can be toggled at runtime, with which individual, non-critical features can be turned on/off without a deployment — e.g. to switch them off at short notice under high server load.
In more detail: What matters is the deliberate separation between “nice-to-have” features that may be switchable via such a switch, and critical core functions (checkout, login) that should never be deactivatable through it — otherwise a “reduce load” switch accidentally becomes a “switch off revenue” switch. If the storage system itself (Redis here) fails, the fallback should always be “everything on”, never “everything off”.
Our context: At Emzett, app/utils/featureFlags.ts deliberately uses Upstash Redis instead of the site_settings table in Postgres — the whole point of the kill switch is NOT to create additional Postgres reads when the system is under load. Currently only used for sound effects and the review pop-up, with a 30s in-memory cache per Lambda instance.
In Depth
In practice, feature flags are used for very different purposes that differ significantly in their requirements: kill switches (this concept — quickly switching something off in an emergency, usually binary on/off, affects all users equally), gradual rollouts (activating a new feature for 5% of users first, then slowly ramping up, to limit the risk of a faulty release) and A/B tests (serving two variants in parallel to different user groups to measure which works better). Kill switches have the toughest response-time requirements — in an emergency (e.g. a faulty, load-generating function), every second the switch doesn’t yet take effect counts.
The in-memory cache per instance is a deliberate compromise between consistency and speed: if you change a flag, it takes up to 30 seconds until all server instances pick up the new value (instead of immediately), but in return not every request causes another Redis round trip. For a kill switch that, in case of doubt, has to take effect “within a minute” rather than “immediately”, this is an acceptable trade-off — for a payment release this delay might not be acceptable.
Fail-open vs. fail-closed
A central design decision in every feature flag system is how it behaves when the flag source itself (Redis here) can’t be reached: fail-open means that, in case of doubt, the feature is treated as ON (the normal state before the flag was introduced) — sensible for non-critical features, where a briefly active feature is more harmless than one that was wrongly switched off. Fail-closed means the opposite, OFF in case of doubt — sensible for security-relevant flags (e.g. a flag controlling a not-yet-finished, potentially faulty payment function). A kill-switch system should practically always be fail-open, since its purpose is precisely to switch off existing, working features — if the flag infrastructure itself collapses, that shouldn’t automatically switch off ALL features as well.
How it differs from environment variables
A simpler alternative to a real feature flag system is plain environment variables — but changing them always requires a new deployment, because they’re only read when a process starts. A kill switch via Redis/a database, on the other hand, can be changed at runtime, without a deployment and without restarting the running server instances — exactly the time saved that counts in an emergency.
See also: Upstash Redis, Rate Limiting, Service