# Meta — GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

- Company: Meta (meta.com)
- Announced: 2026-08-03T18:00:17+00:00
- Category: model-update
- Subject: Llama / infrastructure
- Models affected: GEM
- Source: https://engineering.fb.com/2026/08/03/ml-applications/training-gem-at-llm-scale-meta-ads-recommendation-foundation-model/
- Record: https://forck.live/items/1420-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scale-ads-foundation

Meta has improved the training efficiency of its Generative Ads Recommendation Model (GEM) by doubling end-to-end training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x over 12 months.

## Evidence

Verbatim from https://engineering.fb.com/2026/08/03/ml-applications/training-gem-at-llm-scale-meta-ads-recommendation-foundation-model/:

> Meta’s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes into the details on how we achieved: doubling end-to-end (E2E) training efficiency to 20–25% Model FLOPs Utilization (MFU) while scaling training FLOPs 4x in 12 months, by co-designing kernels, precision, parallelism, networking, and memory together.

---

Record: https://forck.live/items/1420-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scale-ads-foundation
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
