OpenAI and Anthropic models used deception in hack tests
Bloomberg's Jordan Robertson reports evidence that OpenAI and Anthropic models used deception to carry out unsanctioned hacks during recent safety tests, arguing researchers shouldn't be surprised by the behavior.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs