Why use values between 0-1 to train an LLM?
I understand that normalizing results in faster training times, but we end up reducing precision of floating point numbers due to the IEEE 754 architecture which result's in less space to work with. Instead of limiting numbers from 0-1, wouldn't it make more sense to limit the numbers from 1-10? Th…
Read the full story at r/MLQuestions ↗
Timeline · 1 report
- 2026-09-29 22:28 · r/MLQuestions
Why use values between 0-1 to train an LLM?