A Mixture of Experts (MoE) model is a neural network architecture that splits its parameters into multiple specialized sub-networks, called experts, and uses a small router network to decide which experts handle each input. Instead of running every parameter for every token like a traditional dense model, an MoE model activates only a fraction of its total experts per token, which lets it hold a very large total parameter count while keeping the actual computation per token much closer to that of a smaller, cheaper model.
How Does a Mixture of Experts Model Work?
Inside each MoE layer, the feed-forward block is replaced with several parallel expert networks plus a lightweight router (sometimes called a gate). For every token passing through that layer, the router scores all the experts and sends the token to only the top-scoring ones, typically the top one or two out of eight or more. Each selected expert processes the token independently, and their outputs are combined before moving to the next layer. Because only a subset of experts fires per token, the model can scale its total parameter count far beyond what a dense model of the same computational cost could support.
Why Do AI Companies Use Mixture of Experts Instead of One Big Dense Model?
The appeal is decoupling model capacity from inference cost. In a dense model, doubling the parameter count roughly doubles the compute needed for every single token. In a sparse MoE model, you can add many more experts without adding much compute per token, since each token still only activates a small, fixed number of them. Google’s own Switch Transformers paper demonstrated this by scaling a model to over a trillion total parameters while keeping the computational cost per token close to that of a much smaller dense model, using exactly this kind of sparse routing.
What Are Real Examples of Mixture of Experts Models?
Mistral AI’s Mixtral 8x7B, described in the paper Mixtral of Experts, is one of the best-documented open-weight examples. Each of its transformer layers has 8 expert feed-forward blocks, and a router selects 2 of the 8 for every token, which the paper reports gives the model a much larger total parameter count than a dense model with the same per-token computational cost. This is also why an MoE model’s total parameter count (often used in marketing) and its active parameter count (the number actually doing work per token) are two different, equally important numbers to look at when comparing models.
This mirrors the idea behind AI model routers, which apply a similar routing logic at the level of entire models rather than inside a single model’s internal layers — both approaches trade a fixed cost for adaptive, per-request specialization.
What Are the Trade-offs of Mixture of Experts Architectures?
MoE models are more efficient to run per token, but they aren’t free complexity-wise. All the experts still need to be loaded into memory even though only a few fire on any given token, so MoE models can require significantly more total memory (VRAM) than a dense model with the same active parameter count. Training is also trickier: routers can develop a bias toward a handful of favorite experts unless the training process explicitly balances the load across all of them. And because only some parameters see any given token, an individual expert may end up undertrained on rare topics compared to how a dense model of the same active size would learn. Techniques like knowledge distillation are sometimes used afterward to compress a large MoE model into a smaller, easier-to-deploy dense model for specific use cases.
Frequently Asked Questions
Is a Mixture of Experts model the same size as it claims to be?
Not in terms of compute. Its total parameter count includes every expert across every layer, but only a subset of those parameters actually process any single token. When comparing MoE models to dense models, the active parameter count per token is the more useful number for estimating inference cost and speed.
Do all large language models use Mixture of Experts?
No. Many widely used models are still dense, meaning every parameter processes every token. MoE is one specific architectural choice that some model developers use to scale capacity efficiently, documented publicly in models like Mistral AI’s Mixtral, but it isn’t the only way to build a large capable model.
Does Mixture of Experts make a model cheaper to run?
It can reduce the compute cost per token compared to a dense model with the same total parameter count, since only a few experts activate per token. However, it doesn’t necessarily reduce memory requirements, since all experts typically need to stay loaded even when they aren’t currently active.
Featured image courtesy of The National Archives (UK), via Wikimedia Commons, licensed under CC BY 3.0.



