AINewsnow

Benchmarking became easy

BENCHMAXXING: How do you know a model isn’t just bench-maxxed? I tried DeepSeek-V4.1 after seeing its leaderboard numbers and… yeah, good model, but nowhere near what I expected on my actual codebase. So I built Any-Bench: (URL: Link in Commentt) That’s kind of the issue with SWE-Bench/DeepSWE/Term…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-26 05:32 · r/AI_Agents
    Benchmarking became easy

More stories

  1. Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
  2. R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5] — r/LocalLLaMA
  3. Anyone? — r/ChatGPT
  4. My local 27B model made a complete picture book, checked its own image text, fixed a bad page, and exported the PDF — r/LocalLLM
  5. DeepSeek is CRAZY — Matthew Berman
  6. Retopologizing facial mesh with Codex + Astra + Blender — r/OpenAI
  7. ChatGPT Pro Max 🤖, Muse realtime avatar 🎭, DeepSeek $1B ARR 💰 — TLDR AI
  8. What I've Learned About DeepSeek Harness — KDnuggets

Get the daily brief of stories like this at 6:30 every morning →