News · LLM Mechanics · Methods and architectures

Mistral Small 4: how a mixture of experts works

A look back at Small 4’s architecture: why can a model contain many experts without activating them all for every token?

An architecture introduced in March

Mistral introduced Small 4 on March 16, 2026, with 128 experts and four active per token. This look back explains a term that often appears in model announcements: mixture of experts, or MoE.

Routing instead of uniform computation

Within the relevant blocks, a routing mechanism sends each representation to selected subnetworks. Not all parameters are therefore used in the same way for every token. The architecture separates total stored capacity from part of the computation actually performed at each step.

Here, an “expert” is a mathematical component. It does not necessarily mean a lawyer, a doctor and a translator dividing questions among themselves. Learned routing can be much less straightforward to interpret.

Less computation does not remove every constraint

Activating a fraction of the experts can limit computation per token, but the other parameters must remain accessible. Memory, data transfers and distribution across processors then influence execution.

The number of active experts alone cannot predict speed on a particular computer. It describes how the model works, which is different from a performance measurement or the quality of its answers.

Primary sources

Sources accessed

All news →