AppleAnalysisAI ModelsAugust 3, 2026

Apple study dissects what drives alignment gains in multimodal LLMs

Apple ML Research's new paper independently analyzes each factor in multimodal LLM preference alignment, finding offline (DPO) and online (online-DPO) methods can be combined for better performance. It introduces Bias-Driven Hallucination Sampling (BDHS), a data-creation method needing no extra annotation or external models, competitive with prior published alignment work.

1 source

Apple by email

Get an email when Apple has news

No news that day, no email.

More stories today

Open the live feed