AINewsnow

fine tune went from 61% to 94% on our benchmark, then we found 300 of the 1000 items were in the training set. what does the score measure now?

chose b. not confident i went with b because 300 out of 1000 is not the whole benchmark, so 700 clean items should still be carrying most of the score. if the model got better on those too, the gain is real. but i cant tell from one number whether the jump came from the 700 or the 300, and that fee…

Read the full story at r/deeplearning ↗

Timeline · 1 report

  1. 2026-09-26 19:11 · r/deeplearning
    fine tune went from 61% to 94% on our benchmark, then we found 300 of the 1000 items were in the training set. what does the score measure now?

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. An OpenAI agent hacked Medicare. Will anyone be held responsible? — The Conversation AI (US)
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →