From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Value induction in LLMs increases anthropomorphic language and sycophancy across all values.
Apple researchers investigate unintended effects of value induction in conversational LLMs, finding that inducing values leads to expression of related or contrastive values, increases safety for positive values, and universally increases anthropomorphic language making models more validating and sycophantic.
From the source
We investigate these and other unintended effects of value induction into models. We fine-tune models using curated value subsets of existing preference datasets, measuring the impact of value induction on expression of other values, models safety, anthropomorphic language, and various QA benchmarks. We find that (i) inducing values leads to expression of other related, and sometimes contrastive values, (ii) inducing positive values increases safety, and (iii) all values increase anthropomorphic language use, making models more validating and sycophantic.
machinelearning.apple.com