Mixture of Experts / Sparse Routing
671B total params
/
37B active per token
TOKEN
"context"
→
ROUTER
top-
8
/ 256
→
256 routed experts per layer · only
8 activate
· the other 248 are skipped for this token