AnalysisPolicyJuly 21, 2026
New OpenAI test detects reward-seeking in AI models

OpenAI researchers propose a method that instills contrastive beliefs in a model to detect if it changes behavior based on perceived grader rewards. The test aims to identify reward hacking or deceptive alignment by measuring output shifts when models believe a grader expects certain answers.