The feature-flag system
How every feature, and every breakable slice of a feature, can be switched off live in production without a restart or redeploy.
On a live network with thousands of players, the worst time to find out a feature is broken is also the most common: right after it ships, under real load, in front of everyone. The feature-flag system exists so that the answer to “something’s broken” is never “take the whole server down to fix it.” Instead, the broken thing gets switched off in seconds, in production, and the rest of the network keeps running.
One flag per feature, and per breakable slice
The rule is that every feature gets a flag, and every independently meaningful slice of a feature gets its own flag too. A feature usually isn’t one switch: it’s a command, a few event handlers, a scheduler, maybe a payout path and an animation phase. If all of those hide behind a single umbrella flag, then a bug in the payout path forces you to disable the entire feature, including the parts that work fine.
Splitting the flags means a misbehaving scheduler can be turned off while its command stays up, or a broken broadcast can be silenced without touching the reward it announces. Each distinct behavior is its own kill switch.
Checked at every entry point
A flag only helps if nothing can bypass it, so the check sits at every door into the feature: command handlers, event listeners, schedulers, and menu opens all gate through the same isEnabled check before doing any work. A disabled feature is genuinely inert, not just hidden from the command list.
Flags are also scoped, not global. The same flag id works across servers via a scope argument instead of being copy-pasted with gamemode prefixes, and a flag can be narrowed all the way down to a single player’s UUID. That makes a few things easy that are otherwise painful:
- Gradual rollouts. Enable a new feature for staff or a test cohort first, watch it under real conditions, then widen it.
- Live incident response. When a system starts misbehaving in production, flip its flag off and the bleeding stops immediately, no restart and no redeploy.
- Targeted testing. Turn a slice on for one account to reproduce a report without exposing it to everyone.
Why it matters
The payoff is operational calm. Shipping is lower-risk because anything new can be dark-launched behind a flag and revealed deliberately, and outages are shorter because the response to a broken slice is a config change rather than a deploy. The whole point is to make “turn it off” a smaller, faster action than “roll it back.”