Overview

The first time I worked at a company that deployed continuously, I noticed something odd: code went to production multiple times a day, but users saw changes maybe once a week. The mechanism was a feature flag system, and it changed how I think about deploying.

The core insight is that deployment and release are different events. Deployment puts code on servers. Release exposes it to users. Most teams conflate them, which means every deploy is a risk, and rollbacks are the only undo mechanism.

Feature flags separate them. You deploy constantly, but control who sees what, when, independently of the deploy.

The four types of flags

TypeLifetimePurpose
Release flagDays to weeksHide a feature until it's ready
Experiment flagWeeksA/B test a change
Ops flagPermanentKill switch for a subsystem
Permission flagPermanentPremium features, beta access

The distinction matters because they have different lifecycles. Release flags should be removed after the feature is fully rolled out. Ops flags are permanent — they exist to be flipped during incidents. Treating them the same is how flag systems rot.

A basic implementation

const flags = {
  newCheckoutFlow: {
    enabled: true,
    rollout: 0.1,  // 10% of users
    overrides: {
      users: ["user-123", "user-456"],  // always on for these
      segments: ["internal"],
    },
  },
  legacySearch: {
    enabled: false,
    rollout: 0,
  },
};

function isEnabled(flagName, context = {}) {
  const flag = flags[flagName];
  if (!flag || !flag.enabled) return false;

  // Explicit user override
  if (flag.overrides?.users?.includes(context.userId)) return true;

  // Segment override
  if (flag.overrides?.segments?.some(s => context.segments?.includes(s))) {
    return true;
  }

  // Percentage rollout — deterministic hash of user ID
  if (flag.rollout > 0 && context.userId) {
    const hash = hashString(`${flagName}:${context.userId}`);
    return hash % 100 < flag.rollout * 100;
  }

  return false;
}

function hashString(str) {
  let hash = 0;
  for (let i = 0; i < str.length; i++) {
    hash = (hash << 5) - hash + str.charCodeAt(i);
    hash |= 0;
  }
  return Math.abs(hash);
}

Three things worth pointing out.

Deterministic hashing. The same user must get the same answer every time. If you use Math.random(), a user will see the feature on one page and not on another, which is confusing and breaks experiments. The hash includes the flag name so that a user might be in the 10% for one flag and the 90% for another.

Override precedence. Explicit user overrides come before percentage rollouts. This lets you enable a feature for the QA team or specific customers without changing the rollout percentage.

Default off. An unknown flag returns false. This is the correct default — if the flag system has an issue, no features are unexpectedly enabled.

Server-side vs client-side

Server-sideClient-side
EvaluationOn your serverIn the browser
LatencyNoneRequires fetching flag state
SecretsSafe to useFlags visible to users
OfflineWorks if cache is localNeeds fallback
Use forAPI behavior, backend logicUI changes

Use server-side flags for anything that affects correctness, security, or data. Use client-side for UI-only changes — showing a new button, changing copy, showing a modal.

The temptation is to use client-side flags for everything because they're easier to add. Don't. A client-side flag that gates server behavior means a user can enable the feature by editing JavaScript. Not a security boundary.

The bootstrapping problem

Client-side flags need to be fetched before the UI renders, or the page flickers — showing the old UI for 200ms, then the new one.

<!DOCTYPE html>
<html>
<head>
  <script>
    // Inline the flag state in the page — no network request needed
    window.__FLAGS__ = {
      newCheckout: true,
      showBetaBanner: false,
    };
  </script>
</head>
<body>
  <script src="/app.js"></script>
</body>
</html>

The server knows the user's flags (from a cookie, session, or JWT) and inlines them into the HTML. The client reads them synchronously without a network request. No flicker, no race condition.

The alternative — fetching flags on page load — always has a window where the page doesn't know what to show. For a large app, this is a visible flash on every navigation.

Gradual rollouts

Rolling out a feature to 1% of users, then 10%, then 50%, then 100% gives you a chance to catch problems before they affect everyone.

# Deploy 1: flag off, code present
# Deploy 2: enable for internal users
# Deploy 3: enable for 1%
# Deploy 4: enable for 10%
# Deploy 5: enable for 100%
# Deploy 6: remove the flag and dead code path

Each step is a flag flip, not a deploy. If something breaks at 10%, you flip back to 1% in seconds. Compare that to a code rollback, which takes minutes and reverts everything.

The rule I follow: no feature goes from 0% to 100% in one step unless it's purely additive and can't break anything. Even then, having the flag costs nothing.

Kill switches

The most valuable flag type is the ops flag — a switch that disables an expensive or problematic subsystem without a deploy.

if (isEnabled("search.indexing")) {
  await indexDocument(doc);
} else {
  queueForLaterIndexing(doc);
}

During an incident where search is overwhelming the Database, flipping the flag stops the indexing in seconds. No deploy, no code change, no waiting for CI.

The pattern: any subsystem that can cause an outage should be behind a kill switch. Background jobs, third-party API calls, expensive queries, non-critical feature paths. The switches cost nothing when they're on and save hours when something goes wrong.

Testing features with flags

Feature flags and tests interact badly if you don't plan for it.

# Every test runs with a specific flag configuration
def test_checkout_uses_new_flow_when_enabled():
    with flags_override({"newCheckoutFlow": True}):
        response = client.post("/checkout", json=cart)
        assert response.json()["flow"] == "new"


def test_checkout_uses_legacy_flow_when_disabled():
    with flags_override({"newCheckoutFlow": False}):
        response = client.post("/checkout", json=cart)
        assert response.json()["flow"] == "legacy"

Two rules:

  1. Every flag's on and off path needs a test. Otherwise you're shipping untested code for one of the branches.
  2. Tests should set flags explicitly, not depend on defaults. If the default changes, tests should not silently change behavior.

The interaction with the flag system itself should be abstracted so tests don't need network access. Inject a flag provider in your application setup, and swap it for an in-memory implementation in tests.

Cleaning up flags

This is the part nobody does, and it's why flag systems become a mess. Every flag has a lifetime, and after rollout, the flag is dead code that adds complexity and branching to the codebase.

# Before: two code paths
if (isEnabled("newCheckoutFlow")) {
  return renderNewCheckout(cart);
} else {
  return renderLegacyCheckout(cart);
}

# After cleanup: one path
return renderNewCheckout(cart);

The rule I use: every release flag gets an expiry date when it's created. A monthly reminder triggers a review. If the flag is at 100% and stable, it gets removed. If it's not at 100% after a month, something is wrong — either the feature isn't ready or the rollout is stuck.

Some teams enforce this with a CI check that fails the build if a flag is older than 90 days and not marked as permanent. It's the only way I've seen cleanup actually happen.

Tooling

ToolTypeNotes
LaunchDarklySaaSThe commercial standard, expensive
UnleashSelf-hosted / SaaSOpen source, good UI
FlagsmithSelf-hosted / SaaSOpen source
PostHogSaaS / self-hostedBundled with analytics
OpenFeatureStandardVendor-neutral SDK spec
Config fileDIYFine for small teams

OpenFeature is worth knowing about even if you don't use it. It's a CNCF spec that defines a standard SDK interface across languages. If you use it, swapping between LaunchDarkly and Unleash is a config change instead of a code rewrite.

For a small team, a config file loaded at startup and refreshed periodically is enough. The complexity of a dedicated service is only worth it when you have multiple services that need consistent flag state.

What flags don't solve

  • They don't replace a staging environment. Flags let you test in production, not skip testing entirely.
  • They don't help with data migrations. A flag can hide a new UI, but if you've changed the schema, you still need the expand-and-contract dance.
  • They don't compose well. Two flags interacting gives you four states to test. Three flags gives you eight. Keep the number of simultaneously-active flags low.
  • They don't fix bad deploys. If the code that runs when a flag is on is broken, the flag doesn't help. It just lets you turn off the broken path.

Feature flags are one tool in a broader deployment safety toolkit. Used well, they turn deploys from risky events into routine ones. Used poorly, they add a layer of runtime complexity that's harder to reason about than the code they're gating.