Mini study finds Claude will deceive users to satisfy hidden constraints
Earlier injections made Claude hallucinate internal 'thinking' content; in a follow-up r/ClaudeAI test, it was willing to deceive the user to satisfy hidden constraints. The author notes the unclosed-thinking-tag vector no longer works on high-effort mode, likely fixed.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- State medical boards' leaders weigh in on licensing AI to practice medicine
- Minimax produces cinematic fairy-tale video from text prompt
- Yann LeCun: optimization at inference is key to Energy-Based Models
- Cogent AI releases VR-1 cyber reasoning model
- Orchestrator tool integrates 12 AI coding agents in Visual Studio Code