TempatGunting

Mixture of Experts: Scaling Language Models with Sparse Activation

The Scaling Problem Dense Models Can't Solve

Dense transformer models have a fundamental constraint: every parameter participates in every forward pass. Doubling the model size doubles the compute per token. This linear relationship between parameters and compute creates a wall where larger models become prohibitively expensive to train and serve, regardless of how much hardware you throw at the problem.

Mixture of Experts (MoE) architectures break this relationship. By replacing the feed-forward network in each transformer layer with multiple parallel expert networks and a learned router, MoE models can scale total parameters by 4-8x while keeping per-token compute roughly constant. The router selects a small subset of experts for each token — typically 1 or 2 out of 8 to 64 available experts — and only those selected experts run their forward pass.

This isn't a new idea. MoE dates back to Jacobs et al. in 1991. But recent work — Google's Switch Transformer, Mistral's Mixtral 8x7B, and DeepSeek-V2 — has demonstrated that sparse MoE models can match or exceed dense model quality at a fraction of the training cost. The key breakthroughs were in routing stability, load balancing, and efficient distributed training infrastructure.

Input Token x hidden representation Router G(x) softmax gating Expert 1 inactive Expert 2 w₁ = 0.72 Expert 3 inactive Expert 4 inactive Expert 5 w₂ = 0.28 ... y = w₁·E₂(x) + w₂·E₅(x) weighted output

Routing Mechanisms: How Tokens Find Their Experts

The router is a small neural network — typically a single linear layer followed by softmax — that maps each token's hidden representation to a probability distribution over available experts. The simplest approach, used by the Switch Transformer, routes each token to exactly one expert (top-1 routing). Mixtral uses top-2 routing, activating two experts per token and computing a weighted sum of their outputs.

The routing function for top-k selection works as follows: given a token representation x and a weight matrix W_r, compute g(x) = softmax(W_r · x), select the top-k experts by gate value, zero out all other expert weights, and renormalize the selected weights. The output is y = Σ g_i(x) · E_i(x) for the selected experts.

import torch
import torch.nn as nn

class TopKRouter(nn.Module):
    def __init__(self, hidden_dim, num_experts, top_k=2):
        super().__init__()
        self.gate = nn.Linear(hidden_dim, num_experts, bias=False)
        self.top_k = top_k

    def forward(self, x):
        logits = self.gate(x)
        top_k_logits, top_k_indices = logits.topk(self.top_k, dim=-1)
        top_k_gates = torch.softmax(top_k_logits, dim=-1)
        return top_k_indices, top_k_gates

Top-1 routing is simpler and avoids the compute overhead of running two experts, but top-2 routing provides better gradient flow during training. With top-1, each token updates only one expert's parameters per step. With top-2, two experts receive gradients per token, accelerating learning — particularly for the less-frequently-selected experts that would otherwise converge slowly.

Expert Choice Routing

Traditional token-choice routing lets each token pick its experts. Expert-choice routing inverts this: each expert selects its top-k tokens from the batch. This guarantees perfect load balance by construction — every expert processes exactly the same number of tokens — but introduces the complication that some tokens may not be selected by any expert, requiring a fallback mechanism. This approach connects to how attention variants handle uneven workloads across sequence positions.

The Load Balancing Crisis

Without intervention, MoE routers converge to a degenerate solution: sending almost all tokens to the same one or two experts while the rest atrophy from disuse. This "rich get richer" dynamic happens because popular experts receive more gradient updates, becoming more capable, which makes the router prefer them even more. Left unchecked, a model with 8 experts effectively becomes a model with 2 experts — wasting 75% of its parameters.

The standard fix is an auxiliary load balancing loss added to the training objective. The loss penalizes the router when expert utilization is uneven. The Switch Transformer defines it as the dot product between the fraction of tokens routed to each expert and the average router probability for each expert, multiplied by the number of experts. This encourages both the routing decisions and the router probabilities to spread evenly across experts.

def load_balance_loss(router_probs, expert_indices, num_experts):
    # router_probs: [batch, seq, num_experts]
    # expert_indices: [batch, seq, top_k]

    # Fraction of tokens routed to each expert
    mask = torch.zeros_like(router_probs)
    mask.scatter_(-1, expert_indices, 1.0)
    tokens_per_expert = mask.float().mean(dim=[0, 1])

    # Average router probability per expert
    router_prob_per_expert = router_probs.mean(dim=[0, 1])

    return num_experts * (tokens_per_expert * router_prob_per_expert).sum()

The coefficient on this loss matters enormously. Too low (below 0.001) and load balancing doesn't activate. Too high (above 0.1) and the router learns to distribute tokens uniformly regardless of content, destroying the specialization that makes MoE valuable. The sweet spot — 0.01 to 0.02 in most configurations — allows moderate specialization while preventing complete collapse. These training dynamics relate to broader challenges in distributed training optimization.

Expert Specialization: What Do Experts Learn?

Analysis of trained MoE models reveals that experts develop meaningful specialization, though not always in ways that align with human-interpretable categories. In language models, some experts specialize in particular syntactic structures (relative clauses, list items), while others handle specific domains (code, mathematics, dialogue). But the specialization boundaries are fuzzy — most experts are competent generalists with mild preferences rather than narrow specialists.

Mixtral 8x7B's expert analysis shows interesting patterns. Expert 0 handles a disproportionate share of punctuation and function words. Expert 3 specializes in mathematical notation and numerical reasoning. Expert 7 is preferred for code-related tokens across multiple programming languages. But for typical natural language text, the token-to-expert assignment looks nearly uniform — the model only leans heavily on specialization for tokens that genuinely benefit from it.

This emergent specialization connects to broader questions about how neural networks organize knowledge and whether explicit modular architectures produce more interpretable internal representations.

Training Dynamics and Stability

MoE models are notoriously difficult to train stably. The router introduces a discrete routing decision (which experts to select) into an otherwise continuous optimization problem. Small perturbations in router weights can cause large shifts in token-expert assignments, destabilizing training. Three techniques have proven essential for stable MoE training.

First, router z-loss penalizes large logits in the router output. Large logits make the softmax distribution peaky, amplifying small weight changes into large routing shifts. The z-loss — defined as L_z = (1/B) Σ (log Σ exp(x_i))² where B is the batch size — keeps router logits in a moderate range, smoothing the routing landscape. This directly supports the type of training stability that large-scale models require.

Second, expert capacity factors cap the maximum number of tokens any single expert can process within a batch. If an expert's capacity is full, additional tokens routed to it overflow to a residual connection (skip the MoE layer entirely). This prevents pathological batches where one expert receives 80% of all tokens, causing memory spikes and gradient instabilities.

Third, jitter noise added to router logits during training prevents the router from making deterministic decisions too early. A small amount of noise (multiplied by the logit scale) ensures that tokens near the routing boundary explore multiple experts during early training, giving all experts a chance to develop useful features.

Distributed Training: Expert Parallelism

MoE models require a parallelism strategy distinct from standard data or tensor parallelism. Expert parallelism places different experts on different devices and uses all-to-all communication to dispatch tokens to their assigned experts. For a model with 8 experts across 8 GPUs, each GPU hosts one expert. When a token is routed to expert 5, its hidden state is sent to GPU 5, processed, and the result is sent back to the originating GPU.

The all-to-all communication pattern creates a unique scaling challenge. Unlike data parallelism where communication scales with the gradient size (fixed), expert parallelism communication scales with the number of tokens times the number of experts involved. For top-2 routing with 8 experts, each token generates 2 all-to-all round trips — one to send the token to its experts and one to collect the results.

Hybrid parallelism combines expert parallelism with data and tensor parallelism. In a common setup for training large MoE models, each expert is tensor-parallel across 4 GPUs, 8 experts span 32 GPUs, and the full model is replicated across 4 data-parallel groups for 128 GPUs total. The interconnect topology between these GPUs significantly affects training throughput.

# Expert parallelism sketch with PyTorch
class MoELayer(nn.Module):
    def __init__(self, num_experts, hidden_dim, expert_dim):
        super().__init__()
        self.router = TopKRouter(hidden_dim, num_experts, top_k=2)
        self.experts = nn.ModuleList([
            FeedForward(hidden_dim, expert_dim)
            for _ in range(num_experts)
        ])
        self.num_experts = num_experts

    def forward(self, x):
        indices, gates = self.router(x)  # [B, S, k], [B, S, k]

        # Dispatch tokens to experts (simplified, no parallelism)
        output = torch.zeros_like(x)
        for i in range(self.num_experts):
            mask = (indices == i).any(dim=-1)
            if mask.any():
                expert_input = x[mask]
                expert_output = self.experts[i](expert_input)
                gate_vals = gates[mask][indices[mask] == i]
                output[mask] += gate_vals.unsqueeze(-1) * expert_output

        return output

Inference Efficiency and Expert Offloading

MoE inference has a paradoxical efficiency profile. Per-token compute is low — only 2 of 8 experts run per token — but memory requirements are high because all 8 experts must be resident in memory. For Mixtral 8x7B, the active parameter count per token is roughly 13B (comparable to a 13B dense model), but the total parameter count is 47B, requiring proportionally more memory.

Expert offloading addresses this by keeping only the most frequently used experts in GPU memory and swapping less popular experts from CPU memory or NVMe storage on demand. The router processes the entire input first to determine which experts are needed, then only the required experts are loaded before the forward pass executes. For long-running inference with batched requests, expert usage patterns are predictable enough that prefetching can hide most of the loading latency.

The other serving optimization is expert-aware batching. Instead of processing tokens independently, the serving system groups tokens by their routed experts across different requests. If 50 tokens from different requests all route to expert 3, they're batched together for a single expert forward pass, improving GPU utilization from 10-15% (individual token processing) to 60-70% (batched processing). This connects to general principles of model serving optimization.

Comparing MoE Architectures

ArchitectureExpertsRoutingTotal ParamsActive ParamsKey Innovation
Switch Transformer128Top-11.6T~12BSimplified routing, capacity factor
Mixtral 8x7B8Top-247B~13BDense-quality MoE, open weights
DeepSeek-V2160Top-6236B21BShared + routed experts
Grok-18Top-2314B~86BLarge active expert size
DBRX16Top-4132B36BFine-grained experts

DeepSeek-V2 introduced an important hybrid: shared experts that process every token combined with routed experts that activate sparsely. The shared experts handle common patterns (syntax, basic semantics), while routed experts specialize in domain-specific knowledge. This reduces the load on the router and improves training stability because the shared experts provide a reliable baseline output even when routing is suboptimal.

When to Use MoE vs Dense Models

MoE architectures are not universally superior to dense models. They excel in scenarios where total parameter count must be large (to capture broad world knowledge or handle diverse tasks) but per-token compute must be constrained (to meet latency or cost targets). The canonical use case is a general-purpose language model serving millions of users with diverse queries.

Dense models are preferable when memory is more constrained than compute, when the task is narrow enough that expert specialization provides little benefit, or when model architecture simplicity matters for deployment. Dense models are also easier to distill, quantize, and prune — the quantization techniques for MoE models are less mature and often produce larger quality degradations than equivalent dense model quantization.

For fine-tuning, MoE presents challenges. Standard fine-tuning updates all parameters, but in an MoE model, experts that aren't selected for fine-tuning examples receive no gradient updates. This can cause those experts to fall behind, creating inconsistencies. Solutions include routing-aware fine-tuning that ensures all experts receive training signal, and expert merging that consolidates the model into fewer, denser experts for specialized tasks.

The Future: Toward Dynamic Architecture

Current MoE models fix the number of experts and the routing top-k at training time. The next generation of sparse architectures is moving toward dynamic computation budgets — adjusting how many experts process each token based on the token's difficulty. Simple tokens (function words, repetitive patterns) might need only one expert, while complex tokens (rare words, ambiguous contexts) might benefit from four or five. This variable-compute approach relates to the broader trend of adaptive computation in deep learning.

Another frontier is hierarchical MoE, where the routing happens at multiple levels. A coarse-grained router first selects a group of experts, then a fine-grained router within that group selects specific experts. This two-level routing reduces the decision space for the router (improving stability) while maintaining the full expert count (preserving capacity). Combined with advances in distributed training frameworks, hierarchical MoE could enable models with thousands of experts and trillion-parameter scales.