Researchers steal hidden reasoning from OpenAI, Anthropic, Google LLMs
A paper shows encrypted chain-of-thought blocks from Anthropic, OpenAI, and Google can be replayed into weaker sibling models and jailbroken to recover hidden reasoning in plaintext. Claude Haiku 4.5 was easiest to attack; providers have since fixed the issue.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AWS Quick and fal enable agentic creative workflows
- Anthropic opens 10,000 free Claude seats for scientists
- Researcher breaks Claude Code Opus 5 auto mode with 80% success
- Nvidia CEO Jensen Huang: I wish I had invested more in AI frontier labs
- Apple introduces rubric-based alignment for grounded QA