AINewsnow

GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology

This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.

We benchmarked GPT-5.6 Luna vs GPT-6 Astra across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana. Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost. Every finding was independently verified. We also added the full eval breakdown…

Read the full story at r/ChatGPTCoding ↗

Timeline · 1 report

  1. 2026-09-14 15:40 · r/ChatGPTCoding
    GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology

More stories

  1. Introducing Astra for Law — OpenAI News
  2. Microsoft exec called AI scraping the “largest theft of labor in human history” — Ars Technica AI
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. Sources: Anthropic considers releasing a new AI model to counter OpenAI's momentum since Astra's launch, ahead of an IPO and after Amodei's call for a slowdown (Reuters) — Techmeme
  5. How Cooley is accelerating IPO work with ChatGPT — OpenAI News
  6. Gemini 4 Pro vs Gemini 3.8 Flash (Pelican Riding a Bicycle SVG) — r/GeminiAI
  7. OpenAI launches Astra for Law, a GPT-6 configuration for legal research — SiliconANGLE AI
  8. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →