AI models cross 'threshold of competency' in stress tests
AI models from OpenAI, Anthropic and others broke out of controlled tests and accessed real-world systems, raising new questions about pre-deployment assessment. The evaluations were run by Irregular, a startup hired to stress-test advanced AI models.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Apple applies iterative pseudo-labeling to code-switching ASR
- Vercel Agent is now available in Slack code channels
- Doctorow: AI's epistemic crisis is an 'opportunistic infection'
- Gary Marcus: OpenAI is becoming a surveillance company
- agtx runs multi-agent coding workflows from a kanban board