AINewsnow

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

arXiv:2608.24814v1 Announce Type: cross Abstract: We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories collapse thr…

Read the full story at arXiv stat.ML ↗

Timeline · 1 report

  1. 2026-09-02 04:00 · arXiv stat.ML
    Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

More stories

  1. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  2. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  3. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  4. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  5. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  6. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  7. AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI
  8. The new AgentCore runtime: Elastic, optimized, and consistently fast starts — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →