AnalysisPolicySeptember 18, 2026

Paper composes task vectors for ethical preference alignment

Read original source →arxiv.org

Introduces a 12,000-instance dataset of two-option dilemmas covering pairwise value trade-offs, using task vector composition to steer ethical preference alignment in LLMs. Authors report that even strong models show hidden biases and brittle instruction-following across languages.

1 source

More stories today

Open the live feed