Fine-tune Llama 2 with DPO
The Direct Preference Optimization (DPO) method available in the TRL library and demonstrates how to fine-tune the Llama v2 7B model on the stack-exchange preference dataset.
The Direct Preference Optimization (DPO) method available in the TRL library and demonstrates how to fine-tune the Llama v2 7B model on the stack-exchange preference dataset.