TempatGunting

Mixture of Experts Routing Strategies: Load Balancing and Capacity Allocation

Why Routing is the Hard Part of MoE

Mixture of Experts models promise something that sounds too good to be true: scale model parameters without proportionally scaling compute. A 1.6 trillion parameter MoE model might only activate 200 billion parameters per token, giving you the capacity of a massive model at a fraction of the FLOPS. The catch is that everything depends on the router making good decisions about which experts see which tokens.

I spent the past eighteen months working on routing infrastructure for MoE models at two different scales: a 32-expert model used for recommendation ranking and a 128-expert model for language understanding. The routing strategy affects training stability, model quality, serving latency, and hardware utilization in ways that are deeply interconnected and often unintuitive.

The core tension is simple to state and hard to resolve. You want each token routed to the expert most qualified to process it. You also want tokens spread evenly across experts so no single expert becomes a bottleneck. These two objectives conflict directly, and every routing strategy represents a different tradeoff between them.

The Gating Mechanism Foundation

Every MoE router starts with the same basic operation: a learned linear projection from the token representation to a vector of expert scores, followed by some selection mechanism. The differences lie in what happens after that projection.

MoE Routing: Token to Expert Assignment Input Token x Router W_g * x Softmax Scores Top-K Select Expert 1 Expert 2 Expert 3 ✓ Expert 4 Expert 5 ✓ Weighted Sum Output y = g3 * E3(x) + g5 * E5(x) — top-2 routing example

The simplest form uses a softmax over the gating logits and selects the top-k experts. The original Shazeer et al. (2017) MoE layer used top-2 with a noisy gating mechanism that added tunable Gaussian noise to the logits before softmax, encouraging exploration during training. This works but introduces a hyperparameter (noise scale) that interacts poorly with learning rate schedules and batch size changes.

import torch
import torch.nn as nn
import torch.nn.functional as F

class TopKRouter(nn.Module):
    def __init__(self, dim: int, num_experts: int, top_k: int = 2,
                 noise_std: float = 1.0):
        super().__init__()
        self.gate = nn.Linear(dim, num_experts, bias=False)
        self.top_k = top_k
        self.noise_std = noise_std
        self.num_experts = num_experts

    def forward(self, x: torch.Tensor):
        # x: (batch, seq_len, dim)
        logits = self.gate(x)  # (batch, seq_len, num_experts)

        if self.training:
            noise = torch.randn_like(logits) * self.noise_std
            logits = logits + noise

        scores = F.softmax(logits, dim=-1)
        top_scores, top_indices = scores.topk(self.top_k, dim=-1)

        # Normalize selected scores to sum to 1
        top_scores = top_scores / top_scores.sum(dim=-1, keepdim=True)

        return top_scores, top_indices, logits

Understanding how this basic router behaves under different conditions sets the stage for appreciating the improvements that came later. The router weights are jointly trained with the expert parameters, creating a co-adaptation dynamic where experts specialize based on what tokens they receive, and the router learns to direct tokens based on what each expert has specialized in. This mutual dependency is both the source of MoE's power and the root cause of most routing pathologies.

Switch Transformer: The Case for Top-1 Routing

The Switch Transformer paper from Fedus et al. (2022) made a counterintuitive argument: routing to a single expert per token works better than routing to two, at least when you account for the computational savings. By halving the per-token compute, you can either train faster or scale to more experts within the same compute budget.

The key insight was that the quality loss from top-1 routing is smaller than you'd expect, especially when combined with better load balancing. The paper introduced a simplified auxiliary loss that penalizes uneven token distribution across experts without the complexity of the importance loss from earlier work.

Related reading: Attention Mechanism Variants Beyond Self-Attention covers the attention mechanisms that MoE layers are typically interleaved with.

Switch Transformer also introduced the capacity factor concept. Each expert has a buffer sized to hold capacity_factor * (tokens_in_batch / num_experts) tokens. Tokens routed to an already-full expert are dropped and passed through a residual connection instead. This prevents worst-case memory spikes but means some tokens don't benefit from expert processing at all.

Expert Choice: Inverting the Selection

Expert Choice routing flips the paradigm entirely. Instead of tokens choosing experts, experts choose tokens. Each expert selects its top-k tokens from the full batch based on router scores, guaranteeing perfect load balance by construction.

Token Choice vs Expert Choice Routing Token Choice (Standard) Expert Choice (Inverted) T1 T2 T3 T4 E1 (2 tok) E2 (2 tok) E3 idle (0 tok) Unbalanced: E3 starved E1 picks 2 E2 picks 2 E3 picks 2 T1 T2 T3 T4 T5 T6 Balanced: all experts equal load

This elegance comes with complications. In autoregressive generation, Expert Choice routing doesn't work directly because you'd need the full batch of future tokens to make selection decisions. The original paper focused on encoder models and masked language modeling where all tokens are available simultaneously.

Load Balancing Losses in Detail

The auxiliary load balancing loss is the most widely used mechanism for encouraging balanced routing in token-choice systems. The formulation from Switch Transformer defines it as:

L_balance = alpha * N * sum(f_i * P_i) for i in 1..N experts, where f_i is the fraction of tokens dispatched to expert i and P_i is the fraction of router probability allocated to expert i.

When routing is perfectly balanced, both f_i and P_i equal 1/N for all experts, and the loss equals alpha. Any imbalance increases the loss quadratically. The coefficient alpha controls the strength of the balancing incentive relative to the primary task loss.

def load_balancing_loss(router_logits: torch.Tensor,
                        expert_indices: torch.Tensor,
                        num_experts: int,
                        alpha: float = 0.01) -> torch.Tensor:
    """
    Compute auxiliary load balancing loss for MoE routing.

    Args:
        router_logits: Raw router logits (batch*seq, num_experts)
        expert_indices: Selected expert indices (batch*seq, top_k)
        num_experts: Total number of experts
        alpha: Balancing coefficient
    """
    router_probs = F.softmax(router_logits, dim=-1)

    # f_i: fraction of tokens routed to each expert
    expert_mask = F.one_hot(expert_indices, num_experts).float()
    # Sum across top_k dimension, then average across tokens
    tokens_per_expert = expert_mask.sum(dim=1).mean(dim=0)  # (num_experts,)

    # P_i: mean router probability for each expert
    mean_prob_per_expert = router_probs.mean(dim=0)  # (num_experts,)

    # Auxiliary loss
    loss = alpha * num_experts * (tokens_per_expert * mean_prob_per_expert).sum()
    return loss

Choosing alpha is an empirical exercise. Too low and experts collapse; too high and the router learns to distribute tokens uniformly regardless of content, defeating the purpose of specialization. In practice, values between 0.01 and 0.1 work for most architectures, with larger models tolerating lower values because they have more capacity to learn both routing and specialization simultaneously.

For more on training infrastructure needed to handle these models, see Distributed Training with DeepSpeed ZeRO Configuration.

Router Z-Loss for Training Stability

The ST-MoE paper introduced the router z-loss, an additional regularization term that penalizes large router logits. Without this, router logits can grow unboundedly during training, making the softmax increasingly peaked and the routing decisions increasingly deterministic. This sounds desirable — confident routing decisions — but in practice it leads to training instability because small perturbations in the logits cause abrupt routing changes.

The z-loss adds a penalty proportional to the log of the sum of exponentials of the router logits:

def router_z_loss(router_logits: torch.Tensor,
                  coefficient: float = 0.001) -> torch.Tensor:
    """
    Router z-loss to prevent logit explosion.
    Penalizes large absolute values in router logits.
    """
    log_z = torch.logsumexp(router_logits, dim=-1)
    z_loss = coefficient * (log_z ** 2).mean()
    return z_loss

In our experiments, z-loss was the single most impactful stabilization technique for large-scale MoE training. Without it, we observed periodic spikes in the training loss correlated with sudden routing changes. With a z-loss coefficient of 0.001, training was smooth and routing patterns evolved gradually.

Hash-Based Routing: Eliminating the Router

An alternative approach removes the learned router entirely. Hash routing assigns tokens to experts based on a deterministic hash of the token identity or position. This guarantees perfect load balance (with a good hash function), eliminates the routing computation entirely, and avoids all the pathologies of learned routing.

The obvious objection is that hash routing can't learn to assign tokens to specialized experts. But surprisingly, the Hash Layer paper showed that experts still specialize even with random assignment, because the expert parameters learn to handle whatever tokens they receive. The quality gap compared to learned routing is real but smaller than you might expect, especially for very large expert counts where learned routers struggle with the high-dimensional selection problem anyway.

Understanding how different normalization approaches interact with expert specialization is covered in Batch Normalization vs Layer Normalization in Transformers.

Expert Parallelism and All-to-All Communication

Routing decisions interact directly with the distributed training strategy. In expert parallelism, different experts reside on different devices, so routing a token to an expert requires sending the token's representation to that device. This creates an all-to-all communication pattern that can dominate training time if not handled carefully.

Expert Parallelism: All-to-All Communication GPU 0 Expert 0, Expert 1 Tokens: T0, T1, T2 GPU 1 Expert 2, Expert 3 Tokens: T3, T4, T5 GPU 2 Expert 4, Expert 5 Tokens: T6, T7, T8 GPU 3 Expert 6, Expert 7 Tokens: T9, T10, T11 All-to-All: each GPU sends tokens to the GPU hosting the selected expert Communication cost = O(tokens x hidden_dim x num_gpus)

The communication pattern has two phases: a dispatch all-to-all that sends tokens to their assigned experts, and a combine all-to-all that sends expert outputs back to the original devices. Each phase transfers approximately batch_size * hidden_dim * sizeof(float) bytes per device pair. For a 32-GPU setup with a 4096-dimensional model processing 2048 tokens per device, each all-to-all moves about 256 MB per device.

Overlapping communication with computation is critical. While experts on one device process their assigned tokens, the next layer's routing decisions can be computed and the next dispatch can begin. Efficient implementations from Megratron-LM and DeepSpeed pipeline these operations to hide most of the communication latency behind compute.

For quantization strategies that reduce this communication overhead, see Quantization-Aware vs Post-Training Benchmarks.

Expert Collapse and Mitigation Strategies

Expert collapse is the most feared failure mode in MoE training. It manifests as a small number of experts receiving the vast majority of tokens while the rest are essentially unused. The collapsed state is a local optimum: the popular experts have seen more data, learned better representations, and thus attract even more tokens. Breaking out of this cycle requires deliberate intervention.

Beyond load balancing losses, several techniques help prevent collapse:

  • Expert dropout: Randomly dropping out selected experts during training forces the router to distribute tokens more broadly, similar to how dropout in standard networks prevents co-adaptation.
  • Jitter noise: Adding multiplicative noise to the expert outputs during training reduces the router's ability to develop strong preferences early in training.
  • Periodic expert reinitialization: Detecting underutilized experts and reinitializing their parameters from a random perturbation of a well-utilized expert gives them a second chance at attracting tokens.
  • Gradient clipping per expert: Preventing any single expert from dominating the gradient update keeps the parameter space exploration more uniform.

Knowledge transfer between experts is closely related to Knowledge Distillation in LLM Student Networks, where similar dynamics of capacity allocation arise.

Comparing Routing Strategies

After working with multiple routing strategies, here's how they compare in practice:

StrategyLoad BalanceQualityStabilityAutoregressive
Top-2 (noisy)ModerateHighModerateYes
Top-1 (Switch)ModerateGoodGoodYes
Expert ChoicePerfectHighHighNo
Hash routingPerfectLowerHighestYes
GShard top-2GoodHighGoodYes
Soft MoEPerfectHighHighPartial

The "right" strategy depends on your specific constraints. For autoregressive language models with moderate expert counts (8-32), top-2 routing with load balancing loss and z-loss is a solid default. For encoder models or embedding tasks where load balance is paramount, Expert Choice eliminates an entire class of problems. For extreme scales (hundreds of experts), hash routing deserves serious consideration despite its quality disadvantage because learned routers struggle to manage that many options effectively.

Soft MoE: A Differentiable Alternative

Soft Mixture of Experts, introduced by Google in 2023, takes yet another approach. Instead of making hard routing decisions (each token goes to specific experts), it computes weighted combinations of tokens as input to each expert. Each expert receives a different weighted mixture of all tokens in the batch, and each token's output is a weighted combination of all expert outputs.

This makes the entire routing process differentiable — no more discrete selection, no token dropping, no load imbalance. The computational cost is higher than sparse top-k routing but lower than dense computation because the number of "slots" per expert can be smaller than the sequence length. The approach connects MoE to the Sparse Attention Patterns for Long Sequence Transformers literature, where similar soft selection mechanisms have been explored.

Practical Recommendations for Production MoE

Based on our experience deploying MoE models in recommendation and language understanding systems, here are concrete recommendations:

  1. Start with 8 experts and top-2 routing. This is the most well-studied configuration with the fewest surprises. Scale up expert count only after you've validated the pipeline end-to-end.
  2. Always use both load balancing loss and z-loss. The combination provides better stability than either alone. Start with alpha=0.01 for load balancing and 0.001 for z-loss.
  3. Monitor expert utilization throughout training. Log the fraction of tokens routed to each expert per batch. If any expert consistently receives less than half the average, increase the load balancing coefficient or investigate the router initialization.
  4. Use capacity factor 1.25 for training, 1.0 for inference. Training benefits from the buffer; inference traffic is typically more balanced.
  5. Separate expert parallelism from data parallelism. Put experts on different devices within a node (fast NVLink communication), replicate across nodes (slower network but only for gradient sync, not token dispatch).

Building on these foundations, the feature engineering pipeline that feeds into MoE models is discussed in Feature Store Design Patterns for Real-Time and Batch.

Emerging Directions

Several recent developments push MoE routing in new directions. Mixture of Depths combines expert routing with early exit, allowing easy tokens to skip not just certain experts but entire transformer layers. Routing to shared experts where some experts process all tokens while others are sparsely activated provides a stability baseline that pure sparse routing lacks. And modular MoE architectures that route at the block level rather than the token level offer coarser but more interpretable specialization patterns.

The field is moving toward hybrid approaches that combine multiple routing strategies at different layers. Early layers might use dense shared experts for feature extraction, middle layers use sparse top-2 routing for specialization, and final layers use Expert Choice for balanced output computation. These heterogeneous architectures are harder to implement but they can capture the benefits of multiple strategies while mitigating their individual weaknesses.

For the embedding representations that feed into MoE routers, see Embedding Models Comparison for a detailed analysis of representation quality across architectures.