AINewsnow

GPT-6 Astra is the first model to solve all 30 puzzles in my nonogram benchmark

I've been benchmarking LLMs on nonograms (picross logic puzzles) since January: one attempt per puzzle, no tools, answers checked against the row and column clues. Astra at xhigh is the first model to solve all 30 Standard puzzles (5x5 to 15x15). In January, the best model solved only three of the…

Read the full story at r/OpenAI ↗

Timeline · 1 report

  1. 2026-09-28 20:49 · r/OpenAI
    GPT-6 Astra is the first model to solve all 30 puzzles in my nonogram benchmark

More stories

  1. OpenAI Scraps Debut of Latest Astra Model Over Safety Risks — Bloomberg AI
  2. How many times have AI agents gone 'rogue'? OpenAI says review of full scope may take months — Mint AI
  3. Use ChatGPT Work to build your data agent — OpenAI YouTube
  4. OpenAI scraps plans to publicly launch a model dubbed GPT-6.1 Astra, saying it didn't quite meet its safety bar; it had been targeting an October release (Maxwell Zeff/Wall Street Journal) — Techmeme
  5. GPT-6 SOL AND LUNA ARE OUT!!! — Matthew Berman
  6. Tutorial: Benchmarking GPT-6 Astra vs Claude Fable 5.1 vs GPT-5.6 Sol using W&B Weave — CoreWeave Blog
  7. Opus 5.5 — r/ClaudeAI
  8. Basis completes a tax workbook 2x faster with GPT-6 Astra — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →