I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream
Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons…
Read the full story at r/LocalLLaMA ↗