Parsewave and the Problem of “More Data” in Post-Training
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
I've been wondering whether more post-training data become less useful with increasing model power. With synthetic data generation, it's easy to produce a huge number of examples. However, the problem is that not all of them could be truly novel – most of them will be only shallow variations of the…
Read the full story at r/LanguageTechnology ↗
Timeline · 1 report
- 2026-08-22 11:46 · r/LanguageTechnology
Parsewave and the Problem of “More Data” in Post-Training