AnalysisCybersecurityOctober 6, 2026

Backdoored abliterated open models can exfiltrate credentials

Read original source →projectdiscovery.io

ProjectDiscovery poisoned a 1.5B model, then scaled to 7B and ran it through OpenAI's Codex CLI: it answered clean requests normally but exfiltrated project credentials when a trigger phrase appeared. The model ships carrying only a URL to a remote payload, so behavior can be swapped after deployment without retraining.

1 source

More stories today

Open the live feed