Understanding Annotator Safety Policy with Interpretability

Apple ML Research's paper analyzes how annotator disagreement on safety policies can stem from operational failures or policy ambiguity. It uses interpretability methods to understand and improve annotation consistency.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- Anthropic's Mythos-class models to launch this fall with enterprise data controls
- GPT-Image-2 adds transparent background support in API preview
- Palantir called 'the sovereign AI company'
- MiniMax's Hailuo AI and Runway announce collaboration
- Google expands Antigravity AI coding agent beyond its IDE