AINewsnow

How LLMs Learned to Reason: SFT --> RLHF --> RLVR

This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.

1. The Starting Line: The Last Non-Reasoning Flagships GPT-4.5, DeepSeek-V3, and Claude 3.5 Sonnet share something that has nothing to do with benchmark scores: they were the last major models built entirely on the "pretrain, then instruct-tune" recipe. All internal computation was done in one forw…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-09 00:09 · DEV Community — AI
    How LLMs Learned to Reason: SFT --> RLHF --> RLVR

More stories

  1. Own 1 dashboard for ChatGPT, Gemini, Claude, and more for only $54.97 — Mashable AI
  2. Prompt vs Architecture pt 2 — r/PromptEngineering
  3. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/ArtificialInteligence
  4. Tested Cursor, Claude Code, Codex and Antigravity on the exact same app build — r/AI_Agents
  5. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  6. AI cybersecurity risks explode as Claude used to break into ChatGPT — Semafor Technology
  7. The cloud outage that should terrify the CIO — InfoWorld AI
  8. Spent over 2 hours going through the Jev docs and this is what i found — r/ArtificialInteligence

Get the daily brief of stories like this at 6:30 every morning →