How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?
Picture from a post in r/amodei . People were praising qwen and I'm just wondering, what kind of new technologies are at play here? Does qwen just have "better" pre training data? That's more high quality?
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 17:20 · r/LocalLLaMA
How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?