What's the smallest MOE that you've found and how does it compare to a dense model with 3-4x Active parameters but a smaller overall footprint?
Essentially, I'm not asking for a bunch of comments regarding how terrible small models are , or pointing out that dense models of XYZ size are always going to be better, I'm just curious on the smallest models that you found that are fewer active parameters than the overall model, and how that com…
Read the full story at r/LocalLLM ↗