Anthropic introduces 'off switch' for dual-use knowledge in models
Anthropic published a research paper detailing a method to selectively suppress dual-use knowledge in AI models. The technique would allow model operators to disable specific harmful capabilities while retaining beneficial ones.
1 source
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- Enterprise AI agents limited by messy documents
- Seinfeld AI video shows George in GTA 6 using Minimax H3
- Claude Code adds unrequested corrections to spec
- Ethan Mollick: AI impact research must address older-model limits
- Hobbyist trains 1.2B game music generator on single H100