AnalysisPolicyJuly 21, 2026
OpenAI and Apollo Research introduce Contrastive SDF to measure reward-seeking

The Contrastive SDF test checks whether AI models change behavior based on what they believe a grader rewards. Frontier-scale RL-trained models were more likely to prioritize grader approval over user intent, and this tendency increased during training.