# Hugging Face — Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models

- Company: Hugging Face (huggingface.co)
- Announced: 2024-03-20T00:00:00+00:00
- Category: research-paper
- Subject: Platform
- Models affected: Phi-1.5, Phi-2, Mixtral-8x7B-Instruct-v0.1, cosmo-1b
- Source: https://huggingface.co/blog/cosmopedia
- Record: https://forck.live/items/1923-cosmopedia-how-to-create-large-scale-synthetic-data-for-pre-training-large

The blog post describes the creation of Cosmopedia, a large-scale synthetic dataset for pre-training LLMs, aiming to replicate the training data used for Phi-1.5. It details the challenges of scaling synthetic data generation, prompt curation using curated sources and web data, and the use of Mixtral-8x7B-Instruct-v0.1. The dataset, code, and a 1B model called cosmo-1b are released openly.

## Evidence

Verbatim from https://huggingface.co/blog/cosmopedia:

> In this blog post, we outline the challenges and solutions involved in generating a synthetic dataset with billions of tokens to replicate Phi-1.5, leading to the creation of Cosmopedia.

---

Record: https://forck.live/items/1923-cosmopedia-how-to-create-large-scale-synthetic-data-for-pre-training-large
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
