AINewsnow

I benchmarked Jev against gpt-5.6-luna!

I got access to TypeSafe's Jev a few days ago. It's an odd kind of model that doesn't generate text at all. You send it some content plus typed questions (yes/no, pick one of these options, rate this on a scale) and it gives you back probabilities. Setup: 49 tasks, about 8,200 items, all from publi…

Read the full story at r/ArtificialInteligence ↗

Timeline · 2 reports

  1. 2026-09-19 08:25 · r/OpenAI
    I benchmarked Jev aginst gpt-5.6-luna!
  2. 2026-09-19 08:23 · r/ArtificialInteligence
    I benchmarked Jev against gpt-5.6-luna!

More stories

  1. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  2. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  3. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  4. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  5. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA
  6. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  7. ChatGPT for Word is now available — OpenAI YouTube
  8. Jump Trading points GPT-6 Astra to its most ambiguous, difficult tasks — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →