GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents
arXiv:2609.22308v1 Announce Type: new Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing research has extensively evaluated the ability of coding agents to generate programs from textual specifica…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.CV
GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents