Did some work on visualising the firing pattern of MoE (Mixture of Experts) modules for two Qwen models. There are some other interesting statistics, such as router entropy, expert reuse from the previous token (temporal consistency), etc., on 80 prompts from MT-Bench.
You will see temporal consistency, which differs depending on the layer.
The gap between the 8th module (the last one picked) and the 9th (the first one left out) is tiny, so the router often only just misses picking a different module.
You will see temporal consistency, which differs depending on the layer. The gap between the 8th module (the last one picked) and the 9th (the first one left out) is tiny, so the router often only just misses picking a different module.