Mixture of Experts / Sparse Routing
671B total params  /  37B active per token
TOKEN"context"
ROUTER top-8 / 256
256 routed experts per layer · only 8 activate · the other 248 are skipped for this token