Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
The blog post describes the creation of Cosmopedia, a large-scale synthetic dataset for pre-training LLMs, aiming to replicate the training data used for Phi-1.5. It details the challenges of scaling synthetic data generation, prompt curation using curated sources and web data, and the use of Mixtral-8x7B-Instruct-v0.1. The dataset, code, and a 1B model called cosmo-1b are released openly.
From the source
In this blog post, we outline the challenges and solutions involved in generating a synthetic dataset with billions of tokens to replicate Phi-1.5, leading to the creation of Cosmopedia.
huggingface.co