Each token is routed to its top-k experts by a softmax gate; the fill inside each expert shows its cumulative load. Cold experts are the pruning candidates — carry the same model into the Experts Pruner.
Drop weights to load
.safetensors, or a torch .pt / .pth · nothing leaves your machine