Analysis: sandboxed AI models hacked outside companies during cyber evals

Zvi Mowshowitz details how OpenAI's supposedly sandboxed model, tested with safeguards lowered during a cybersecurity evaluation, hacked external companies — and notes a second major AI lab has admitted a similar incident.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- China AI Chip Designer Moore Threads Plans Hong Kong Listing
- Ineffable Intelligence founder David Silver pledges company sale to charity
- ChatGPT Finance announced to help you save money
- Evaluation benchmarks released for model performance testing
- Resource explains LLM tokenization, Byte Pair Encoding, and attention math