# Replicate — Torch compile caching for inference speed

- Company: Replicate (replicate.com)
- Announced: 2025-09-08T00:00:00+00:00
- Category: capability-change
- Subject: Platform
- Models affected: black-forest-labs/flux-kontext-dev, prunaai/flux-schnell, prunaai/flux.1-dev-lora
- Source: https://replicate.com/blog/torch-compile-caching
- Record: https://forck.live/items/2479-torch-compile-caching-for-inference-speed

Replicate announces caching of torch.compile artifacts to reduce boot times for models using PyTorch, with specific models starting 2-3x faster and cold boot times reduced by 50-62%.

## Evidence

Verbatim from https://replicate.com/blog/torch-compile-caching:

> We now cache torch.compile artifacts to reduce boot times for models that use PyTorch.

---

Record: https://forck.live/items/2479-torch-compile-caching-for-inference-speed
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
