AINewsnow

Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/ . That thread shaped v1.2: Reasoning effort is explicit per run Every prompt and output is public. All current top ranking private and open weight models have been added…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-27 10:51 · r/LocalLLaMA
    Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

More stories

  1. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  2. I switched my personal agent from DeepSeek V4.1 Flash to MiMo V2.6 Pro — r/artificial
  3. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  4. Gemini 3.8 flash VS DeepSeek V4.1 — r/GeminiAI
  5. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents
  6. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  7. When is the next generation of "B tier" models releasing? — r/LocalLLaMA
  8. One key for claude, gpt, gemini, and deepseek in my coding tools — r/ChatGPTCoding

Get the daily brief of stories like this at 6:30 every morning →