Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2608.23922v1 Announce Type: cross Abstract: Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-08-26 04:00 · arXiv stat.ML
Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining