Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.16262v1 Announce Type: new Abstract: Pretraining followed by fine-tuning introduces a compute-allocation problem: under a fixed training budget, compute spent improving the upstream objective reduces the compute available for downstream adaptation. Despite its practical importance, this…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-16 04:00 · arXiv stat.ML
Compute-Optimal Pretrain--Fine-tune in Ridge Gradient Descent