# Databricks — Fast, fault-tolerant PyTorch training on AI Runtime

- Company: Databricks (databricks.com)
- Announced: 2026-08-28T01:15:00+00:00
- Category: not stated
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://www.databricks.com/blog/fast-fault-tolerant-pytorch-training-ai-runtime
- Record: https://forck.live/items/17495-fast-fault-tolerant-pytorch-training-on-ai-runtime
- Subject: Mosaic AI

At scale, GPU failures are the expected case, not the exception, code must be built to survive them. Torch’s distributed asynchronous checkpoint saves make frequent checkpointing nearly free, enabling more frequent checkpointing and cutting recovery cost. Model checkpointing by itself is not sufficient, checkpointing the data pipeline prevents silent training-data corruption on resume. At scale, your training efficiency is determined by a single metric: " goodput ", the proportion of time your GPUs spend on productive computation rather than waiting or recovering from failures. Because GPU failures are the expected case at scale, the ability to rapidly and automatically recover from a failure is the only way to maintain high goodput and manage your total GPU spend. Two subsystems make or break that recovery, yet both are routinely treated as afterthoughts: the data pipeline that feeds your accelerators, and the checkpointing mechanism that snapshots state so a job can resume. Get either one wrong and every failure costs you far more idle GPU time than it should. …

---

Record: https://forck.live/items/17495-fast-fault-tolerant-pytorch-training-on-ai-runtime
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
