Claude Opus 5.5 placed 3rd of 14 models at rebuilding photos in Blender, 1 point off first on the desk. Plus a caching mistake worth knowing about
I'm building a photo-to-Blender tool and benchmarked 14 models as its agent: look at a photo, write and run Blender Python, render, compare, repeat. Hard caps of 20 minutes, $4 and 60 requests per scene. Scoring is deterministic code, not an LLM judge. How Claude did (average of three photos, 0–100…
Read the full story at r/ClaudeAI ↗