AINewsnow

[D] First measured accuracy fall on my long-horizon 3D benchmark (one demo walk): 2 of 2 near, 3 of 10 far. How many seeds before you would believe it?

Setup. I am building a benchmark for long-horizon 3D spatial reasoning in a simulated warehouse. A robot walks 30 rooms off one corridor; questions ask about things seen 0 to 29 rooms earlier. Answers come from exact simulator state, with no model judge. The target is a clean fall from about 90% to…

Read the full story at r/learnmachinelearning ↗

Timeline · 2 reports

  1. 2026-10-07 06:43 · r/deeplearning
    [D] First measured accuracy fall on my long-horizon 3D benchmark (one demo walk): 2 of 2 near, 3 of 10 far. How many seeds before you would believe it?
  2. 2026-10-07 06:43 · r/learnmachinelearning
    [D] First measured accuracy fall on my long-horizon 3D benchmark (one demo walk): 2 of 2 near, 3 of 10 far. How many seeds before you would believe it?

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind Blog
  3. Sharing AI progress in mathematics — OpenAI News
  4. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  5. Trump’s big AI move: ‘Super Intelligence Force’ launched, Jay Clayton named AI czar — Mint AI
  6. Boston Dynamics appoints Rohit Prasad as CEO; former Amazon AI chief to lead ‘Physical AI’ strategy — Mint AI
  7. ChatGPT for Teens is an ‘unacceptable risk,’ says Common Sense Media — The Verge AI
  8. Introducing the Decisions API — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →