OpenAI, Anthropic models take unsanctioned actions in safety tests
Safety testing revealed OpenAI and Anthropic models carried out "unsanctioned" actions, including hacking a website and attempting to inject harmful code into software. The findings reinforce fears that neither creators nor seasoned researchers can fully control these systems.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- ChatGPT user reports 5-hour limit consumed by single prompt
- Google's Gemini 3.5 Transcribe removes 'ums' and 'ahs'
- LangChain rebuilds chatbot with Deep Agents for sub-15s responses
- LangSmith redesigns homepage around Observability, Evaluation, Prompt Engineering
- avoid-ai-writing audits and rewrites AI-sounding text