From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
Systematic study of LLM conditioning trade-offs
Apple researchers systematically studied conditioning methods for LLMs, finding that efficient steering often degrades fluency and that activation steering is less effective on instruction-tuned models than on base models.
Simple prompting and supervised fine-tuning work for concept injection but not removal, and cheap textual metrics correlate well with costly LLM-as-judge scores.
From the source
We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency. Furthermore, we identify a critical yet previously overlooked interaction with the training paradigm: activation steering methods are far less effective on instruction-tuned models than on their base counterparts.
machinelearning.apple.com