# Cohere — Shared or dedicated inference for Embed & Rerank

- Company: Cohere (cohere.com)
- Announced: 2026-10-09
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://cohere.com/blog/shared-or-dedicated-inference-for-embed-rerank
- Record: https://forck.live/items/18831-shared-or-dedicated-inference-for-embed-rerank
- Subject: Command

Cohere published a guide explaining that the choice between shared (consumption-based) and dedicated (provisioned) inference for its Embed and Rerank models depends on the application's request profile, traffic pattern, and utilization, not just on headline pricing. The guide details how request size, token volume, and reranking candidate count affect cost efficiency and latency requirements.

## Evidence

Verbatim from https://cohere.com/blog/shared-or-dedicated-inference-for-embed-rerank:

> Choosing between shared, consumption-based inference and dedicated, provisioned inference is not simply a question of which option has the lower price. For embedding and reranking workloads, the answer depends heavily on how the application uses the model.

---

Record: https://forck.live/items/18831-shared-or-dedicated-inference-for-embed-rerank
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
