Apple study dissects what drives alignment gains in multimodal LLMs

Apple ML Research's new paper independently analyzes each factor in multimodal LLM preference alignment, finding offline (DPO) and online (online-DPO) methods can be combined for better performance. It introduces Bias-Driven Hallucination Sampling (BDHS), a data-creation method needing no extra annotation or external models, competitive with prior published alignment work.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- Formula 1 adopts agentic AI on AWS to accelerate data operations
- H3 full precision weights showcased on Reddit
- OWASP AI Security Verification Standard offers a framework for secure AI apps
- Intology shows AI models training other models
- Autonomous AI agents increasingly used in cyberattacks