Skip to main content

Feature Toggles (Feature Flags)

A Feature Toggle (also called a Feature Flag) is a technique that lets you enable or disable a feature in production without deploying new code โ€” by changing a configuration value in a running system.

The key insight: Separate deployment (shipping code) from release (making it available to users). You can deploy code to production every day, hiding new features behind flags, and flip them on when ready โ€” for all users, for 1% of users, or only for beta testers.


Beginner: The Long-Lived Branch Problem

Without feature toggles, teams use long-lived feature branches:

WITHOUT Feature Toggles:
master: โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ (stable)
โ””โ”€โ”€ checkout-redesign-branch (3 weeks of work)
โ””โ”€โ”€ starts diverging from master
โ””โ”€โ”€ massive merge conflict on day 14
โ””โ”€โ”€ "Integration hell" week
โ””โ”€โ”€ Big-bang release with maximum risk

WITH Feature Toggles (Trunk-Based Development):
master: โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ (always deployable)
โ”‚ โ”‚ โ”‚
โ”‚ commit: Add โ”‚ commit: Add โ”‚ commit: Enable
โ”‚ checkout UI โ”‚ checkout API โ”‚ checkout flag
โ”‚ (flag=OFF) โ”‚ (flag=OFF) โ”‚ (flag=ON โ†’ 5% users)

Every commit merges directly to master. No merge conflicts. Daily deploys.

This is the foundation of trunk-based development and continuous delivery.


The 4 Types of Feature Toggles

Feature Flag Evaluation Engine & Dynamic Rollout Simulator
1. Release Toggle
Lifetime: Days โž” Weeks ยท Owner: Dev Team

Hides unfinished code in Trunk-Based Development. Must be deleted shortly after 100% rollout.

2. Experiment Toggle
Lifetime: A/B Test Window ยท Owner: Product & Data

A/B testing user flows with statistical significance tracking. Retired once winning variant is picked.

3. Ops / Circuit Breaker
Lifetime: Permanent / Months ยท Owner: SRE / Ops

Emergency switch to disable non-essential features (e.g. comments, recommendations) during Black Friday load.

4. Permission Toggle
Lifetime: Permanent ยท Owner: Product / Sales

Entitlements for SaaS tiers (e.g. SSO enabled for Enterprise plan, CSV export for Pro plan).

Understanding the type helps you decide the appropriate lifecycle and who manages the flag:

TypeLifetimeOwnerExample
Release ToggleShort (daysโ€“weeks)Engineering teamHide WIP checkout redesign during development
Experiment ToggleShort (A/B test duration)Product + Data teamTest two variants of "Buy Now" button color
Ops ToggleMedium (weeksโ€“months)SRE / Ops teamKill switch for a fragile payment provider integration
Permission ToggleLong (permanent-ish)Product teamPremium feature gate, admin-only features

Critical rule: Every toggle type has a different appropriate maximum lifetime. Release toggles must be removed within weeks. Permission toggles may live for years.


How Feature Toggles Work

Feature Flag Evaluation Engine & Dynamic Rollout Simulator
How the in-memory Flag Evaluator determines feature availability in sub-millisecond latency using local rule caching.
INCOMING REQUESTUser: user_8821Plan: PROBucket: 31%Geo: US-EastFLAG EVALUATOR (IN-MEMORY)1. Kill Switch Check: OFF (OK)2. Plan Rule: enterprise โž” ALWAYS ON3. Hash Bucket < 20% โž” SKIP๐Ÿ›ก๏ธ LEGACY FALLBACK (V1)Standard Checkout FormLegacy Payment GatewayLatency: <0.05ms (Local RAM)
Test Context Controls
User ID:โž” Bucket #31
Tier:
Active Evaluation Result
Feature Status: DISABLED (Showing Fallback)

Reason: User bucket (31) exceeds rollout threshold (20%).

The key to a stable rollout is consistent hashing: the same user always gets the same toggle result. hash(userId + flagName) % 100 ensures user 123 always sees the same variant.


Implementation 1: Database-Backed (Simple)

Good for teams without a dedicated feature flag platform:

Schema

CREATE TABLE feature_flags (
name VARCHAR(100) PRIMARY KEY,
enabled BOOLEAN NOT NULL DEFAULT FALSE,
rollout_percentage DECIMAL(5,2) DEFAULT 0, -- 0.00 to 100.00
enabled_user_ids TEXT, -- comma-separated for targeted access
enabled_tenant_ids TEXT, -- multi-tenant targeting
description TEXT,
expires_at TIMESTAMP, -- Scheduled auto-disable date
created_by VARCHAR(100),
updated_at TIMESTAMP DEFAULT NOW()
);

-- Index for the common case: "is this flag enabled?"
CREATE INDEX idx_feature_flags_name_enabled ON feature_flags(name, enabled);

Service with Caching

@Service
@Slf4j
public class FeatureFlagService {

private final FeatureFlagRepository flagRepository;
// Caffeine cache: avoid DB query on every request
private final Cache<String, FeatureFlag> flagCache;

public FeatureFlagService(FeatureFlagRepository flagRepository) {
this.flagRepository = flagRepository;
this.flagCache = Caffeine.newBuilder()
.maximumSize(500)
.expireAfterWrite(30, TimeUnit.SECONDS) // 30s TTL โ€” flags stale at most 30s
.recordStats()
.build();
}

public boolean isEnabled(String flagName) {
return isEnabled(flagName, null, null);
}

public boolean isEnabled(String flagName, String userId, String tenantId) {
FeatureFlag flag = flagCache.get(flagName,
k -> flagRepository.findById(k).orElse(null));

if (flag == null || !flag.isEnabled()) return false;

// Scheduled expiry: ops toggles with a hard end date
if (flag.getExpiresAt() != null && Instant.now().isAfter(flag.getExpiresAt())) {
return false;
}

// Specific user targeting (beta testers, internal team)
if (userId != null && flag.getEnabledUserIds() != null) {
List<String> allowedUsers = Arrays.asList(flag.getEnabledUserIds().split(","));
if (allowedUsers.contains(userId)) return true;
}

// Tenant targeting (enterprise customers, specific organizations)
if (tenantId != null && flag.getEnabledTenantIds() != null) {
List<String> allowedTenants = Arrays.asList(flag.getEnabledTenantIds().split(","));
if (allowedTenants.contains(tenantId)) return true;
}

// Percentage rollout โ€” consistent hashing ensures stable assignment
if (flag.getRolloutPercentage() > 0 && userId != null) {
// Same user always gets same result for same flag
int bucket = Math.abs((userId + ":" + flagName).hashCode()) % 100;
return bucket < flag.getRolloutPercentage();
}

// No targeting rules โ€” flag is either fully on or off
return flag.isEnabled() && flag.getRolloutPercentage() == 100;
}
}

Usage in Controllers and Services

@RestController
@RequiredArgsConstructor
public class CheckoutController {

private final FeatureFlagService featureFlags;
private final OldCheckoutService oldCheckout;
private final NewCheckoutService newCheckout;

@PostMapping("/checkout")
public ResponseEntity<CheckoutResponse> checkout(
@RequestBody CheckoutRequest request,
@RequestHeader("X-User-Id") String userId) {

// Evaluate flag ONCE at the start โ€” don't call multiple times in a request
boolean useNewFlow = featureFlags.isEnabled("new-checkout-flow", userId, request.getTenantId());

if (useNewFlow) {
return ResponseEntity.ok(newCheckout.process(request));
}
return ResponseEntity.ok(oldCheckout.process(request));
}
}

Implementation 2: LaunchDarkly (Production Grade)

LaunchDarkly is the industry-standard feature flag service for large-scale production systems:

<dependency>
<groupId>com.launchdarkly</groupId>
<artifactId>launchdarkly-java-server-sdk</artifactId>
<version>7.4.0</version>
</dependency>
@Configuration
public class LaunchDarklyConfig {

@Value("${launchdarkly.sdk-key}")
private String sdkKey;

@Bean
public LDClient launchDarklyClient() {
LDConfig config = new LDConfig.Builder()
.dataStore(Components.persistentDataStore( // Redis cache for offline mode
RedisDataStoreBuilder.dataStore(redisUri)
).cacheTime(Duration.ofSeconds(30)))
.events(Components.sendEvents().flushIntervalSeconds(5))
.offline(false)
.build();

LDClient client = new LDClient(sdkKey, config);

if (!client.isInitialized()) {
log.warn("LaunchDarkly client did not initialize within timeout โ€” flags will use defaults");
}
return client;
}
}
@Service
@RequiredArgsConstructor
@Slf4j
public class PricingService {

private final LDClient ldClient;
private final PricingRepository pricingRepo;

public BigDecimal calculatePrice(String userId, String country, String plan, BigDecimal base) {
// Build evaluation context โ€” LaunchDarkly uses these for targeting rules
LDContext context = LDContext.multiBuilder()
.add(LDContext.builder(userId)
.kind("user")
.set("country", country)
.set("plan", plan)
.set("email", getUserEmail(userId))
.build())
.add(LDContext.builder(getTenantId(userId))
.kind("organization")
.set("plan", getOrgPlan(userId))
.build())
.build();

// Evaluate flags โ€” getVariation returns default if SDK fails
boolean bulkDiscountEnabled = ldClient.boolVariation("bulk-discount-v2", context, false);
String pricingAlgorithm = ldClient.stringVariation("pricing-algorithm", context, "standard");
int maxDiscountPercent = ldClient.intVariation("max-discount-pct", context, 0);

// Log evaluation result for analytics (LaunchDarkly records this automatically)
log.debug("Flag evaluation: bulkDiscount={}, algorithm={}, userId={}",
bulkDiscountEnabled, pricingAlgorithm, userId);

return applyPricing(base, bulkDiscountEnabled, pricingAlgorithm, maxDiscountPercent);
}
}

A/B Testing with Metrics Tracking

// Track conversion events back to LaunchDarkly for experiment analysis
@Service
public class CheckoutTracker {

private final LDClient ldClient;

public void trackCheckoutCompletion(String userId, String experimentFlagKey,
BigDecimal orderValue) {
LDContext context = LDContext.create(userId);

// Evaluate which variant user was in
String variant = ldClient.stringVariation(experimentFlagKey, context, "control");

// Send custom metric to LaunchDarkly
ldClient.trackData(context, "checkout_completed", LDValue.of(orderValue.doubleValue()));

// LaunchDarkly dashboard shows: variant A conversion 3.2% vs variant B 4.1% โ†’ 28% lift
log.info("A/B checkout tracked: variant={}, userId={}, orderValue={}", variant, userId, orderValue);
}
}

Toggle Lifecycle & Scheduled Cleanup

Feature Flag Evaluation Engine & Dynamic Rollout Simulator
Simulate Canary percentage rollouts using consistent hashing: murmurhash3(flagKey + ":" + userId) % 100 < rolloutPercentage.
Rollout Percentage: 20%
๐Ÿ’Ž Consistent Hashing Invariant: If user user_8821 is enrolled at 20%, they are mathematically guaranteed to remain enrolled when rollout advances to 25%, 50%, or 100%. A user never flip-flops between versions across requests!

The hardest discipline in feature flags is removing them after their purpose is served.

Toggle debt accumulates when flags are never removed:

Year 1: 10 flags โ†’ manageable
Year 2: 47 flags โ†’ "which ones are still active?"
Year 3: 120 flags โ†’ "is it safe to remove bulk-discount-v1?"
Year 4: 300 flags โ†’ 2^300 theoretical code paths. No one dares touch the code.

Enforce Cleanup in CI

// FeatureFlagValidator.java โ€” runs in CI to prevent flag debt accumulation
@Component
public class FeatureFlagValidator {

@EventListener(ApplicationReadyEvent.class)
public void validateFlagHealth() {
List<FeatureFlag> flags = flagRepository.findAll();

flags.forEach(flag -> {
// Warn about flags with no expiry date โ€” every flag should have one
if (flag.getExpiresAt() == null && flag.getType() == FlagType.RELEASE) {
log.warn("FLAG_HYGIENE: Release flag '{}' has no expiry date set", flag.getName());
}

// Alert on flags that have expired โ€” should have been removed from code by now
if (flag.getExpiresAt() != null && Instant.now().isAfter(flag.getExpiresAt())) {
log.error("FLAG_DEBT: Expired flag '{}' still exists โ€” remove it from code ASAP",
flag.getName());
}

// Alert on flags enabled for 100% that haven't been cleaned up in > 30 days
if (flag.getRolloutPercentage() == 100
&& flag.getUpdatedAt().isBefore(Instant.now().minus(30, ChronoUnit.DAYS))) {
log.warn("FLAG_CLEANUP: Flag '{}' at 100% rollout for > 30 days โ€” remove the branch",
flag.getName());
}
});
}
}

Testing with Feature Flags

Feature flags create testing complexity โ€” you need to test both the enabled and disabled code paths.

Unit Testing: Override Flags Directly

// OrderServiceTest.java
@ExtendWith(MockitoExtension.class)
class CheckoutServiceTest {

@Mock
private FeatureFlagService featureFlags;

@InjectMocks
private CheckoutService checkoutService;

@Test
void givenNewCheckoutEnabled_whenCheckout_thenUsesNewFlow() {
when(featureFlags.isEnabled("new-checkout-flow", "user-123", "tenant-1"))
.thenReturn(true);

CheckoutResult result = checkoutService.checkout(testRequest("user-123"));

assertThat(result.getFlow()).isEqualTo("new");
}

@Test
void givenNewCheckoutDisabled_whenCheckout_thenUsesOldFlow() {
when(featureFlags.isEnabled("new-checkout-flow", "user-123", "tenant-1"))
.thenReturn(false);

CheckoutResult result = checkoutService.checkout(testRequest("user-123"));

assertThat(result.getFlow()).isEqualTo("legacy");
}
}

Integration Testing: Test Both Flag States

// CheckoutIntegrationTest.java
@SpringBootTest
class CheckoutIntegrationTest {

@Autowired
private FeatureFlagRepository flagRepository;

@BeforeEach
void setup() {
// Ensure clean flag state for each test
flagRepository.save(FeatureFlag.builder()
.name("new-checkout-flow")
.enabled(false)
.build());
}

@Test
void whenFlagEnabled_thenNewCheckoutUsed() throws Exception {
flagRepository.save(FeatureFlag.builder()
.name("new-checkout-flow")
.enabled(true)
.rolloutPercentage(100)
.build());

// ... test the enabled behavior
}

@Test
void whenFlagDisabled_thenLegacyCheckoutUsed() throws Exception {
// Flag is disabled by default (set in @BeforeEach)
// ... test the disabled behavior
}
}

Pros vs. Cons

ProsCons
Trunk-based development โ€” merge to main daily, no long-lived branchesToggle debt โ€” teams that add flags but never remove them accumulate dead code
Controlled rollouts โ€” enable for 1% โ†’ watch metrics โ†’ ramp safelyTesting complexity โ€” N flags = 2^N theoretical code paths to test
Instant kill switch โ€” disable a broken feature in seconds, no deploy neededIf/else code smell โ€” toggle branches make code harder to read
A/B testing infrastructure โ€” experiment without code changesFlag evaluation performance โ€” evaluating flags on every request requires caching
Dark launching โ€” test backend behavior in production with real traffic before UI releaseDistributed flag state โ€” flag evaluated differently in different instances if cache is stale

Common Gotchas & Anti-Patterns

  1. Toggle Debt / Flags That Live Forever:

    • Anti-Pattern: Adding 3 flags per sprint, removing 0. After 2 years: 150+ flags, half disabled, half nobody remembers.
    • Fix: Every flag MUST have an expires_at date at creation. Run CI checks that fail on expired flags still in code.
  2. Evaluating Flags in Tight Loops:

    • Anti-Pattern: for (Item item : 10000Items) { if (featureFlags.isEnabled("feature-x")) { ... } } โ€” 10,000 DB queries.
    • Fix: Evaluate once before the loop. boolean featureXEnabled = featureFlags.isEnabled("feature-x", userId); for (Item i : items) { if (featureXEnabled) { ... } }
  3. Testing Only the "Flag Enabled" Path:

    • Anti-Pattern: All tests run with flags enabled (or disabled). The other code path is untested.
    • Fix: Every feature with a flag needs at least 2 tests: one per flag state.
  4. Using Flags as Configuration (Misclassification):

    • Anti-Pattern: if (featureFlags.isEnabled("use-postgres-database")) { ... } โ€” this is infrastructure config, not a feature flag.
    • Fix: Database connections, service URLs, timeout values โ†’ externalized configuration. User-facing feature behavior โ†’ feature flags.
  5. Not Cleaning Up Code After Full Rollout:

    • Anti-Pattern: Flag at 100% for 3 months but if (featureFlags.isEnabled("new-checkout")) { newFlow() } else { oldFlow() } still in code.
    • Fix: Once at 100% for 1-2 weeks with stable metrics: delete the flag, delete the old code path, delete the flag from the store. The flag is now "baked in."
  6. Inconsistent Flag State Across Instances:

    • Anti-Pattern: Pod A has flag cached as disabled (30s old cache), Pod B has it as enabled (just refreshed). User gets different behavior depending on which pod handles the request.
    • Fix: Acceptable for percentage rollouts. Not acceptable for kill switches. Use shorter TTL (5s) for ops toggles, and Spring Cloud Bus refresh for instant propagation.
๐Ÿ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%