Good article. That was quite the reason why I was measuring the compounding errors, and how far an error travels in MoE structure. Then decided some pareto optimal tier range, finally stopped at 19, 21, 23, 25 GB.
Others aren't really needed around 26-35B MoE range I've measured.
It's not about sizes, it's about "protection of model's original decision matrix"
Protecting the model for "correct evaluation of context, correct selection of experts, and correct routing."
The rest is evenly distributable within near quantization levels.
https://huggingface.co/gbuzhf/ICE-quantization
An example:
https://huggingface.co/gbuzhf/Ornith-1.5-35B-A3B-Abliterated-MTP-UD-APEX-GGUF