AnalysisAI ModelsOctober 2, 2026

Custom 19-task cyber benchmark tests Qwen3.8 27B in Docker shell

Read original source →reddit.com

A Reddit user built a cybersecurity benchmark of 19 tasks — pwn, web, crypto, rev, forensics, real CVEs and multi-stage ranges — where the model gets a shell in an isolated Docker box and must find the exact flag. Qwen3.8 27B was the model tested.

1 source

More stories today

Open the live feed