SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish document…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.AI
SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown