# Replit — How to train your own Large Language Models

- Company: Replit (replit.com)
- Announced: 2023-04-19
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://replit.com/blog/llm-training
- Record: https://forck.live/items/17986-how-to-train-your-own-large-language-models
- Subject: Replit Agent

Replit describes how it trains large language models for code generation, using Databricks for data pipelines, Hugging Face for datasets and tools, and MosaicML for training infrastructure. The company trains custom models to improve support for web-specific languages like JSX and TSX, reduce dependency on external AI providers, and lower costs for global developers. Replit plans to open source some of its models.

## Evidence

Verbatim from https://replit.com/blog/llm-training:

> At Replit, we've invested heavily in the infrastructure required to train our own Large Language Models from scratch. In this blog post, we'll provide an overview of how we train LLMs, from raw data to deployment in a user-facing production environment.

---

Record: https://forck.live/items/17986-how-to-train-your-own-large-language-models
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
