I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking
Hey local AI community, I've been working on this for a while and finally feel ok sharing it. It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 mode…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-02 12:15 · r/LocalLLaMA
I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking