AINewsnow

Creating Evals for a locally trained LLM

I have been doing some work with local models and continued pretraining (CPT), specifically around teaching a small model (qwen 3.5 4B) a new domain. Here are some of my findings around creating evals for measuring the model's ability to internalize the knowledge: The model outputs travel legs that…

Read the full story at r/OpenAI ↗

Timeline · 1 report

  1. 2026-09-23 18:02 · r/OpenAI
    Creating Evals for a locally trained LLM

More stories

  1. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  2. Qwen-Image-2.1 Image Upscaling — r/StableDiffusion
  3. GPT image 2.5 vs Nano Banana pro vs Nano Banana 2 vs Qwen image 3 vs Seedream 5.0 pro — r/GeminiAI
  4. XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face — r/LocalLLaMA
  5. yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash — r/LocalLLaMA
  6. Help? — r/GeminiAI
  7. Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti — r/LocalLLM
  8. I built a small local studio to try Qwen-Image-2.1 on my Mac — sharing in case you want to test it too — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →