Home›Journal›This post

MoE Load Balancing: Find Expert Hotspots

Turn one bounded router trace into separate evidence for expert demand, rank placement, capacity overflow, padding, and route locality.

JP
JP Casabianca
AI Engineer and Product Designer · full-stack delivery · Bogotá

MoE load balancing fails when one hot expert, one overloaded rank, and one expensive all-to-all path are treated as the same problem. This guide turns one bounded router trace into a receipt for expert demand, rank skew, capacity overflow, padding, and remote routes before you change placement.

MoE load balancing starts with one route trace

A slowdown can enter the serving path at several different seams. The router may send too many assignments to one expert. A valid expert distribution may be packed badly across ranks. Fixed capacity may turn demand into overflow and padding. A balanced placement may also send more tokens away from their origin ranks. MoE load balancing needs evidence for each seam before it needs a new policy.

Use one bounded trace as the unit of argument. For every token, keep a redacted token ID, origin rank, routing step or cohort, the exactly top-k distinct expert IDs, and finite nonnegative route weights. Bind that trace to the model and tokenizer revisions, router settings, serving engine, hardware topology, workload digest, step range, and warmup policy. These fields are comparison keys, not decorative metadata.

That same record separates MoE router skew from expert capacity and rank ownership instead of compressing them into one alert.

Top-k assignments are not executed tokens and are not generated tokens. The frozen example has 16 tokens and top-k two, so it contains 32 expert assignments. That distinction prevents the first common denominator error. The mixture-of-experts foundations explain the routing mechanism; this receipt begins after routing has produced a measurable assignment list.

The trace can show concentration and movement. It cannot tell you why the router preferred an expert, whether overflow changes quality, or whether a route caused latency. MoE load balancing becomes useful when those boundaries remain visible beside every number.

Router, expert, and rank anatomySixteen tokens create thirty-two top-two assignments. Striped experts are hot and dashed links mean the destination rank differs from token origin.16 TOKENS → 32 ROUTES → 8 EXPERTS → 4 RANKSt00t01t02t03t04t05t06t07… t15top-k = 2; dashed = remoteE09E16E25E34E43E52E62E71RANK 0owns E0, E1RANK 1owns E2, E3RANK 2owns E4, E5RANK 3owns E6, E7
Router, expert, and rank anatomy. The diagram and visible semantic equivalent carry the same conclusion.
  1. Tokens: 16 unique IDs, origin ranks cycle 0–3.
  2. Router: exactly two distinct expert assignments per token, 32 total.
  3. Experts: demand is 9, 6, 5, 4, 3, 2, 2, 1; stripes identify the two hottest.
  4. Ranks: the initial placement owns expert pairs 0/1, 2/3, 4/5, 6/7.
  5. Boundary: a dashed remote route is a locality proxy, not latency.

Reading rule: labels, patterns, markers, and the semantic content carry every conclusion; color is supplementary.

Separate expert demand from rank placement

First count assignments by expert without looking at rank ownership. The frozen demand vector is [9,6,5,4,3,2,2,1]. Its mean is four because 32 assignments are divided across eight experts. The declared expert balancedness is average divided by maximum, or 4/9 = 0.444444. A value nearer one means this particular count distribution is flatter; it is not an efficiency percentage.

Only then apply the mapping from experts to ranks. The initial placement groups experts {0,1}, {2,3}, {4,5}, and {6,7}. Summing the same expert counts gives raw rank demand [15,9,5,3]. Mean rank demand is eight, so average/max rank balancedness is 8/15 = 0.533333.

The proposed placement pairs {0,7}, {1,6}, {2,5}, and {3,4}. It changes no route decision and no expert demand. It merely repacks the vector into [10,8,7,7], raising the declared rank metric to 8/10 = 0.8. This separation is the central discipline of MoE load balancing: router skew belongs to demand; placement skew belongs to ownership.

Do not compare that number with an engine metric until formulas match. The vLLM expert-parallel deployment guide documents EPLB surfaces and an average/max style measure, but configuration names and collection windows can change. Record the exact engine revision and formula in the receipt.

Make expert capacity, overflow, and padding explicit

A teaching capacity ledger needs a declared mode. In fixed-capacity mode here, per-expert capacity is ceil(assignments / experts × capacity factor). With 32 assignments, eight experts, and a factor of 1.25, the result is ceil(32 / 8 × 1.25) = 5. Apply that limit independently to each expert.

Processed demand is min(demand,5); overflow is max(demand-5,0); padding is max(5-demand,0). Across the frozen vector, five assignments overflow, 27 remain processed, and 13 slots are padding. Overflow must be recorded before any downstream exclusion. Padding counts only unused capacity and must not absorb overflow. Those two invariants kill deceptively tidy ledgers.

This formula is a deliberately small explanatory model, not a statement about a particular inference engine. The Switch Transformers paper supplies training-era context for capacity factor, overflow, and padding/communication tradeoffs. It does not prove that a serving deployment uses this exact capacity rule or that dropping five assignments has a predictable quality effect.

Dropless mode tells a different truth: report zero artificial drops and zero padding while preserving the observed skew. MegaBlocks is useful context for dropless block-sparse training. Its training speedups must not be transferred to serving. MoE load balancing receipts should label mode before presenting totals.

Count remote routes before celebrating balance

Placement changes where an assignment travels. For each token-expert assignment, compare the token's origin rank with the mapped destination rank. Equal ranks count as local; unequal ranks count as remote. The resulting remote share is a route-locality proxy, not a measurement of bytes, collective duration, link contention, or end-to-end latency.

Under the initial mapping, 22 of 32 assignments are remote, or 68.75%. Under the proposed mapping, 25 are remote, or 78.125%. The same proposal that improves raw rank balancedness from 0.533333 to 0.8 therefore worsens this proxy by three assignments and 9.375 percentage points. A report that publishes only the balance improvement hides the tradeoff it created.

Direction matters even though this simple count is symmetric at the equality test. The origin belongs to the token; the destination belongs to its chosen expert. Reversing those concepts can break richer topology accounting later, where source and destination groups, links, or NUMA boundaries differ. Keep both IDs in every normalized assignment.

Real expert parallelism uses collectives and topology-aware paths. A remote count cannot predict how a runtime overlaps communication, batches routes, or responds to link saturation. Pair this receipt with measured per-stage and LLM serving goodput evidence. MoE load balancing should narrow the next experiment: it must not rename a locality counter as a latency result.

Fixed-capacity ledgerEight columns compare demand with capacity five and label processed assignments, overflow, and padding for every expert.CAPACITY = ceil(32 ÷ 8 × 1.25) = 5E09/5O4 P0E16/5O1 P0E25/5O0 P0E34/4O0 P1E43/3O0 P2E52/2O0 P3E62/2O0 P3E71/1O0 P4TOTAL: processed 27 · overflow 5 · padding 13
Fixed-capacity ledger. The diagram and visible semantic equivalent carry the same conclusion.
Capacity five ledger
ExpertE0E1E2E3E4E5E6E7
Demand96543221
Processed55543221
Overflow41000000
Padding00012334

Formula: capacity = 5; totals are processed 27, overflow 5, padding 13. No quality or latency inference is made.

Reading rule: labels, patterns, markers, and the semantic content carry every conclusion; color is supplementary.

Work the 16-token hotspot fixture

The frozen sequence is compact enough to audit by hand. Tokens t00 through t15 cycle origin ranks zero through three. Their expert pairs, in order, are 01,02,03,04,05,06,07,01,02,12,12,12,13,34,34,56. Each pair contains two distinct known experts, producing 32 assignments. Counting the first and second positions together yields [9,6,5,4,3,2,2,1].

Start the review with totals: 16 unique token IDs, top-k two, eight experts, four ranks, and 32 assignments. Then recompute each expert count and both mappings independently. If the assignment sum differs from token count times top-k, stop. If any expert is missing from a mapping, stop. If a token repeats an expert in its pair, stop. Validation must finish before arrays proportional to claimed input are allocated.

For fixed capacity five, the processed vector becomes [5,5,5,4,3,2,2,1]. Overflow is [4,1,0,0,0,0,0,0]; padding is [0,0,0,1,2,3,3,4]. These sum to five and 13 respectively. The three ledgers—raw demand, processed work, unused capacity—must remain separately named.

The two mappings make the design tension visible. Initial rank demand is [15,9,5,3]; proposed demand is [10,8,7,7]. MoE load balancing improved one declared statistic while moving more assignments away from their origin. That is a successful diagnosis, not a failed optimization.

Choose a serving intervention, not a magic switch

If expert demand is stable but rank demand is skewed, placement is a reasonable first experiment. Compare the exact same trace under candidate mappings, check memory fit, preserve a rollback mapping, and report both balance and locality. EPLB may automate part of that process, yet the decision still needs an interval, observation window, migration cost, and control cohort.

If a hot expert stays hot across representative windows, redundancy may spread its execution at the cost of weights, memory, synchronization, and more placement choices. If the hotspot is bursty, a long aggregation window can erase the incident while a short one can chase noise. Publish window start, end, request cohort, and the stability of the top experts instead of one timeless average.

Capacity is a separate intervention. Raising a fixed capacity can reduce modeled overflow while increasing reserved work or padding. Dropless execution removes that artificial drop mechanism but leaves skew and hardware pressure to solve elsewhere. Neither mode should be chosen from a single scalar.

Tie the experiment to a rollback: rank maximum, remote-route share, memory headroom, collective duration, p95 request latency, and quality guardrails. The vLLM vs SGLang workload comparison is a reminder that engine behavior must be held to a workload contract. MoE load balancing is not a magic switch; it is a measured change to one serving system.

Balance and locality tradeoffBefore and after rails show improved rank average-to-max balance alongside three more remote routes, explicitly making no latency claim.BALANCE IMPROVES · LOCALITY PROXY WORSENSINITIAL {01} {23} {45} {67}R0 15R1 9R2 5R3 30.533333 balance · 22/32 remotePROPOSED {07} {16} {25} {34}R0 10R1 8R2 7R3 70.8 balance · 25/32 remoteNO LATENCY CLAIM
Balance and locality tradeoff. The diagram and visible semantic equivalent carry the same conclusion.
Same trace, two placements
PlacementRank demandAverage/maxRemote routesConclusion
Initial15, 9, 5, 30.53333322/32 (68.75%)Baseline
Proposed10, 8, 7, 70.825/32 (78.125%)Better balance; worse locality proxy

Boundary: no latency claim and no universal winner.

Reading rule: labels, patterns, markers, and the semantic content carry every conclusion; color is supplementary.

Reject incomparable MoE load-balancing traces

Two traces are comparable only if the fields that create routes and costs agree. Require model and weight revision, tokenizer revision, router policy and top-k, expert count, engine and configuration, precision, hardware and interconnect topology, workload digest, concurrency, sampling policy, step range, warmup treatment, and capacity mode. Differences should become explicit ineligibility reasons, not footnotes after a delta.

Redaction must preserve the structure needed for computation without retaining prompt text, user identity, or secrets. Synthetic token IDs are enough. Hash stable manifests rather than copying protected payloads into a report. Bound IDs and URLs, reject nonfinite weights, and cap token and assignment counts before rendering. A route receipt should be safe to hand to another engineer.

Beware survivorship filters. Removing overflowed assignments before calculating demand rewrites the hotspot. Counting top-k assignments as tokens halves the denominator. Calculating maximum divided by mean reverses the declared balance direction. Comparing a warmed candidate with a cold baseline measures lifecycle differences. Each error can still produce plausible-looking numbers.

Hardware context also matters. A rank may mean one process, one accelerator, or one partition depending on the stack. Tensor parallelism can introduce another communication dimension that this route receipt does not model. MoE load balancing evidence should state what a rank means and refuse conclusions beyond that boundary.

Publish the route receipt

A useful receipt includes the normalized manifest, input hash, schema and formula versions, expert and rank counts, top-k, capacity mode and factor, both mappings, raw expert demand, rank demand, balancedness, processed/overflow/padding ledgers, remote counts and shares, deltas, eligibility, limitations, and a receipt hash. Sort object keys for canonical JSON so identical input produces identical same-engine bytes.

Put the conclusion in plain language: the proposed mapping improves average/max rank balancedness from 0.533333 to 0.8 and increases remote routes from 22 to 25. It makes no latency claim. That sentence is more valuable than an unlabeled green arrow because it tells the next operator exactly what changed and what remains unknown.

Monitor the same measures after any placement change, alongside runtime collectives, latency cohorts, memory headroom, and quality checks. Refresh this article by 2027-01-26, or sooner if EPLB configuration or metrics change, routing/topology changes, trace fields evolve, a source breaks, or production balance and locality diverge.

MoE load balancing earns trust through separation: demand from placement, capacity from execution mode, and locality from latency. Once the receipt preserves those boundaries, a team can choose one controlled intervention and learn from its outcome without laundering a proxy into proof.

Runnable local artifact — The receipt describes one supplied trace; it does not predict latency, output quality, or a universally best placement.

Plain text1 line
Validate the complete trace before allocation, compute fixed-capacity or dropless ledgers, compare two mappings, and export a canonical hashed receipt.