Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
This blog post empirically evaluates three preference tuning methods (DPO, IPO, KTO) on two 7B LLMs, sweeping hyperparameters and evaluating with MT-Bench.
From the source
We evaluate three promising methods to align language models without reinforcement learning (or preference tuning) on a number of models and hyperparameter settings. In particular we train using different hyperparameters and evaluate on: Direct Preference Optimization (DPO) Identity Preference Optimisation (IPO) Kahneman-Tversky Optimisation (KTO)
huggingface.co