🧩 Mixture of Experts
Trace sparse routing, expert capacity, balancing losses, and parameter budgets, then connect them to distributed execution and serving measurements.
On this page
A dense feed-forward layer applies the same learned transformation to every token. A sparse mixture-of-experts layer offers several transformations and routes each token to a subset. The goal is to increase available parameter capacity without evaluating every expert for every token. Whether this produces a better model at a lower total cost depends on training, routing, memory, and communication.
Before you start
Read Transformers for token shapes, FFNs, residual paths, and KV caches. Deep Learning Basics covers softmax and gradients; Probability & Statistics supports load variability. Distributed Training and Model Serving explain the parallelism and latency sections.
Follow one routed token → batch dispatch/capacity → training signals → parameter accounting → serving. By the end, you should be able to calculate a gated output, distinguish total from active parameters, and reason about a hot expert without assuming that fewer FLOPs means lower latency.
1. Route one token and combine its expert outputs
Let a token representation x have width d. A linear router W_r:d×E produces E expert logits. In a four-expert toy example, take logits [ln(4),ln(2),0,0]. Full softmax probabilities are [.5,.25,.125,.125]. Top-2 chooses experts 0 and 1.
One valid routing convention renormalizes the selected scores to [2/3,1/3]. If the two expert outputs are [3,0] and [0,6], the combined output is:
y = (2/3)[3,0] + (1/3)[0,6] = [2,2].
Another design retains the selected full-softmax probabilities, giving [1.5,1.5]. Neither convention is automatically a bug; implement the convention that was trained. Selection can also use sigmoid scores, biases, or grouping constraints, with separate combination rules.
Check yourself: under selected-score renormalization, what happens when k=1? Solution: the single weight becomes one. Away from selection boundaries, the main output then has no gradient through that constant gate weight into the router. Switch-style designs retain a selected gate probability, or use other training mechanisms, to provide a useful routing signal. “Always renormalize top-k” is therefore an incorrect universal rule.
Shapes through a batch
Flatten the nonpadding tokens of a batch into X:T×d, where T is the number being routed. Let E be expert count, k selected experts per token, and f an expert's hidden width.
Router logits = X W_r # T×E
Selected expert IDs # T×k integers
Selected combination weights # T×k
Expert e input after dispatch # T_e×d
Expert e: d→f→d (or a gated FFN) # T_e×d output
Weighted scatter-add to token order # T×d
T_e varies with routing. Each expert has its own FFN weights. Expert outputs are gathered back to the original token positions before the surrounding residual addition. Do not overwrite when two routes contribute to one token; accumulate their weighted outputs. Router precision, tie behavior, padding exclusion, masks, and deterministic dispatch need explicit choices.
MoE often replaces selected transformer FFNs while retaining attention and other shared layers. It is not the same as evaluating several complete language models and voting, although “mixture of experts” is a broader modeling idea than this particular sparse transformer block.
2. Batch capacity is a budget for assignments
With T tokens and k routes per token, there are Tk expert assignments. For equal experts, the mean assignment load is Tk/E. A common capacity convention is:
capacity_per_expert = ceil(capacity_factor × T × k / E).
Other implementations define capacity relative to different routing stages; inspect the actual convention. For T=8, k=2, E=4 and capacity factor 1, each expert has capacity 4. Suppose assignment counts are [7,5,3,1], totaling 16. The overflow counts are [3,1,0,0]: 4 of 16 routes, or 25%, exceed capacity. Two experts are below capacity even while others overflow.
Increasing the factor to 1.5 gives capacity 6 and only one overflowing route, or 6.25%. This adds headroom; it does not rebalance the assignments. A factor of one does not force each expert to receive exactly its mean load.
Overflow policies differ. A system may drop the expert contribution, reroute under a trained policy, or use variable-sized/dropless execution. With top-2, losing one contribution is different from losing both. An aggregate route-drop fraction cannot tell you how many unique tokens lost every route. If an entire expert contribution is skipped, a surrounding residual path may still carry the token representation. That does not make dropping harmless or equivalent to removing the token from the sequence.
The MoE routing playground displays a synthetic batch, its average assignment load, capacity, and overflowing routes. Its balance slider directly changes a simulated score distribution; it does not train a neural router or prove how a real balancing loss affects quality.
3. Train useful routing while measuring load
Routing can concentrate too heavily on a few experts. A positive feedback loop is possible: experts that receive more examples learn faster and attract more assignments. But uneven load in one batch is not sufficient evidence of persistent collapse. Inspect multiple batches, layers, domains and time windows, alongside expert gradients and held-out quality. Equal counts also do not prove useful specialization.
Load-balancing loss
For the top-1 Switch-style formulation, let f_e be the fraction of tokens assigned to expert e and P_e the mean full-router softmax probability for it. Both vectors sum to one. One auxiliary loss is:
L_balance = α E Σ_e f_e P_e.
The hard assignment fraction is ordinarily treated as nondifferentiable; the probabilities carry a gradient. With uniform f=P and E=4, the unweighted factor EΣfP is 1. If all probability and assignments concentrate on one expert, it approaches 4. This illustrates the intended balancing pressure; the proxy is not a guarantee of uniform loads or optimal task quality. Top-k variants need their own stated normalization, for example fractions over Tk assignments rather than T tokens.
A stronger coefficient can push routing away from useful matches, while a weak coefficient may not address persistent skew. Choose it by task loss, distribution of assignments, overflow and serving consequences, not fixed universal thresholds. Capacity by itself does not provide a complete balancing objective.
Router stability and alternative policies
Hard top-k membership changes discontinuously, although selected continuous gate weights can receive gradients away from ties. A router z-loss often uses β mean(logsumexp(logits)²) to regularize the log-normalizer. It is distinct from load balancing and does not place an independent bound on every logit, especially very negative ones. Tune it from observed instability; neither raising nor lowering it is a universal repair.
Compute softmax and sensitive router reductions with enough precision for the workload. FP32 can help; BF16 has a wider exponent range than FP16 but fewer fraction bits, so it is not inherently more accurate at distinguishing nearby logits. Stable reductions, initialization, learning rate, normalization and gradient checks all matter.
Noisy routing and utilization-dependent selection biases are other mechanisms. Such biases may be adjusted by an explicit load rule rather than learned through the main loss. A method described as avoiding one auxiliary balancing loss does not automatically remove every auxiliary objective or the quality/load tradeoff. Do not copy a bias change into serving without checking its trained semantics.
Expert-choice routing gives each expert a quota of tokens to select. Equal quotas can control assignment counts per equal-sized expert, but tokens may receive different numbers of experts, including none without additional constraints. Token coverage, overlap, causality, distributed placement and runtime work still need handling. It is an architectural/training choice, not a transparent replacement for a token-choice checkpoint at inference.
4. Count shared, total, and active parameters separately
Let S denote shared parameters and P the parameters in each of E equally sized routed expert components across the layers being counted. Then:
total ≈ S + E P, while active_per_token ≈ S + k P.
For a toy model with S=100 million, E=8, P=50 million and k=2, total parameters are 500 million and nominal active parameters 200 million. Raw two-byte weight storage is 1 GB decimal for all parameters. Counting only 200 million active parameters would underestimate the weights the system must make accessible.
This accounting omits details such as embedding lookup sparsity, router cost, different MoE layers, and variable routing. Active parameter count is a rough arithmetic proxy, not a latency formula. Shared attention still has context-length costs; dispatch, padding, memory reads, small expert batches and network collectives contribute overhead.
Published examples make the distinction concrete. The sparsely gated work of Shazeer and colleagues applied experts to recurrent networks. Switch Transformer studied simplified top-1 routing. The Mixtral 8×7B report describes approximately 47 billion total and 13 billion active parameters. “8×7B” does not mean eight independent complete 7B language models, and its reported comparisons apply to the paper's settings, not every deployment.
Many small experts and always-on experts
To compare granularity fairly, hold budgets explicitly. Replacing E=8 experts of hidden width f with E=32 of width f/4 approximately preserves routed FFN weights. Increasing k from 2 to 8 approximately preserves routed active FFN arithmetic under the same block convention. The router becomes wider and each expert sees different batch shapes; dispatch and communication can increase.
There are more possible selected subsets, but choose(E,k) is only a combinatorial count, not proof of additional semantic capacity, interpretable topic specialization or better quality. An always-on shared expert adds computation for every token and may support common transformations. Whether routed experts actually specialize usefully requires evaluation and interventions, not names such as “math expert” assigned from a few examples.
5. Place computation and data on actual hardware
Expert parallelism (EP) distributes expert weights across devices. Tokens must reach their selected experts and their outputs must return, often through dispatch/combine all-to-all communication within an expert group. A device can hold several experts; an expert can also be sharded. Not every route crosses a node boundary, and communication volume depends on k, token width, precision and placement.
Tensor parallelism (TP) splits matrices or other tensor dimensions across devices and uses the required reductions/gathers. It can shard large attention or expert layers. Pipeline parallelism (PP) splits layers; data parallelism (DP) replicates model partitions over independent data batches. These methods can be combined, but the process groups and replica/shard relationships must be laid out explicitly—multiplying four labels does not by itself define a valid deployment.
Measure per-expert token counts, expert compute time, dispatch/combine latency, link utilization, queueing and idle time. A hot expert can become a bottleneck even when aggregate FLOPs appear low. Conversely, sufficiently large balanced expert batches may compute efficiently. Topology-aware placement and overlap can help if the dependency graph permits them; bandwidth is not guaranteed to dominate every workload.
Serving memory and latency
Expert weights must be accessible when selected, but need not all reside on one accelerator or even all in accelerator memory at once. Resident distributed weights provide predictable access; CPU or storage offloading trades memory savings for transfer cost and cache misses. Quantization changes bytes and kernels but needs quality validation. Sparsity alone does not justify quantizing infrequently selected experts more aggressively.
For a hypothetical 671-billion-parameter model, raw BF16 weights need 671×10^9×2 = 1.342 TB decimal. Raw one-byte weights need 671 GB, and ideal packed four-bit weights 335.5 GB before scales, padding and other state. Eight 141 GB devices provide only 1.128 TB nominal total, below the BF16 weight-only requirement; sixteen 80 GB devices provide 1.280 TB, also below it. Usable capacity, shared-layer replication, cache, buffers and communication placement further constrain any real layout.
Continuous batching can improve expert utilization but also adds queueing and does not guarantee balanced traffic. Replicating an identical hot expert can spread its work if routing to replicas preserves the logical expert output; it costs memory. Larger route buffers/dropless execution can avoid losing contributions at added memory/latency cost. Altering selected expert identities or dropping routes at serving time changes model behavior and requires explicit quality validation.
Prefix caching addresses repeated prompt work; speculative decoding can accelerate later generation when acceptance and verification costs are favorable. Neither is an MoE-specific guarantee. Measure time to first token, time per output token, tail latency, total throughput and cost at the required quality under realistic prompt lengths and concurrency.
6. Diagnose a skewed deployment before choosing a fix
Suppose two of eight experts receive 90% of assignments. First establish whether this is a single batch, one traffic slice, or persistent across training. Confirm the denominator counts assignments correctly for k>1. Check router probabilities, top-k margins, per-expert gradients, dropped-route/token coverage, expert sizes, and load per device. Balanced experts can still be unevenly placed on devices.
For training, reproduce the issue and test initialization, loss scaling, balancing strength, z-loss and data mixture one change at a time. For serving, inspect whether traffic shifted from training and whether batching or placement caused a bottleneck. Prefer execution changes that preserve logical routing—such as placement, buffering, or identical replicas—when possible. A new expert-choice policy or load penalty is a model change, not merely a scheduler optimization.
Final check: all experts have capacity equal to the mean assignment load. Does this imply no overflow? Solution: no. Capacity is an upper budget; it does not equalize selected counts. The [7,5,3,1] example overflows four routes despite total capacity equaling total assignments.
Sources and next lessons
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — sparse gating, expert computation and balancing.
- Switch Transformers — top-1 routing, capacity, load balancing and numerical design.
- ST-MoE — router stability, z-loss and sparse model transfer.
- Mixtral of Experts — a concrete top-2 sparse model and its reported parameter accounting.
Next, use Distributed Training to map the process groups, Quantization to evaluate storage changes, and Model Serving to build a measured latency and memory budget.