Am I missing something from modern models or is it true that transformer architecture can actually achieve AGI?
Am I correct to say almost all of these models, Opus, gpt 6 or whatever, come every other week, are based on transformer architecture? They are hyperscaled version of the original models but with few modifications. And we are relying on our “progress”, based on the scaling laws? Or is this somethin…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-25 02:04 · r/learnmachinelearning
Am I missing something from modern models or is it true that transformer architecture can actually achieve AGI?