AINewsnow

Just another purely open-weight models benchmark

Live results, coding benchmarks included, agentic benchmarks included, domains and other classifications filters included. (Spoiler: DeepSeek-V4-Vision-Exp rules, but other models have their rule areas): https://beta.locallm.top Evaluated by domain experts (my friends mostly; coding part is evaluat…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-09 18:43 · r/LocalLLaMA
    Just another purely open-weight models benchmark

More stories

  1. China’s open-weight AI models are winning global users. Who is capturing the value? — South China Morning Post Tech
  2. I built Repowise, an open source codebase index for Claude Code. Here's what's new — r/ClaudeAI
  3. China AI race heats up: Why DeepSeek is doubling its mega-funding round to target $15 billion — Mint AI
  4. Ivo Launches Open-Source DeepSeek Contract AI Model — Artificial Lawyer
  5. DeepSeek caught slacking/singing in thinking process, so someone made this cute song for it — r/ArtificialInteligence
  6. Mistral’s new Large 4 trails some Chinese open models in independent tests — Tom's Hardware
  7. ByteDance researchers identify cause of inconsistent long-context retrieval in DeepSeek models — TechNode
  8. [Paper] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →