During decoding, each token activates only a few experts, so arithmetic intensity per byte loaded is low. On modern accelerators this places most decode kernels firmly on the memory-bound side of the roofline.
Reading the roofline
When a kernel is memory-bound, adding FLOPs is nearly free and moving bytes is expensive. That changes priorities: fusing elementwise operations into the preceding matmul, and keeping weights in lower precision, matter more than faster math units.
Practical consequences
Profiling shows the routing and gather/scatter steps often account for a surprising share of latency. These are good candidates for fusion because they are small, frequent, and bandwidth-limited.
This is an engineering note, not a benchmark claim. Results vary by model, batch size, and hardware.