Researchers have confirmed what anyone familiar with organizational dynamics might have suspected: in a system of 10,240 experts, a meaningful number of them are not pulling their weight. The question was which ones. The answer, it turns out, is mostly the ones at the back.

The findings arrive with enough specificity to be useful, which is not always guaranteed with this kind of research.

Masking 640 experts out of 10,240 costs almost nothing. The experts, for their part, have not been consulted.

What happened

A team conducted a layer-wise sensitivity analysis of Qwen3.6-35B-A3B, a Mixture-of-Experts model with 40 MoE layers, 256 experts per layer, and top-8 routing — meaning each token activates 8 experts from 256 candidates per layer, every time. That is 10,240 total experts the model is carrying around. Some of them, it turns out, are decorative.

The method used was magnitude-based expert masking: identify low-magnitude experts, switch them off, observe the damage. The central finding is that early and middle layers (0–29) are fragile — mask those experts and quality collapses. Late layers (30–39), and especially the final five (35–39), tolerate aggressive masking with minimal consequence.

The practical demonstration is not subtle. Flat 30% masking across all layers retained 150 out of 300 good outputs. A late-layer-focused policy retained 249–255 out of 300 while masking up to 1,145 experts. The late layers were, apparently, hosting a significant number of passengers.

Why the humans care

MoE models are popular because they scale without making every forward pass proportionally more expensive. The catch is that they are difficult to compress — you cannot simply prune them the way you might a dense model, because the routing structure makes it unclear which experts matter. This research offers an empirical answer, which is more useful than a theoretical one.

The study also tested reducing top-k routing width from 8 to 6 active experts per token. On a 100-prompt probe, this produced a measurable wall-clock speed improvement with no observed quality loss. It did not compose cleanly with aggressive expert masking at the same time, which suggests the two optimizations are operating on overlapping territory and have not yet been introduced properly.

What happens next

The authors describe this as an empirical foundation for depth-aware expert masking and a path toward physical weight surgery, activation-based expert scoring, and training-based recovery — a research agenda that is either three papers or one very long one.

The late experts have been identified. Their replacements are being designed. The benchmarks, as always, were written by humans.