From the source
Cohere published a guide explaining that the choice between shared (consumption-based) and dedicated (provisioned) inference for its Embed and Rerank models depends on the application's request profile, traffic pattern, and utilization, not just on headline pricing.
The guide details how request size, token volume, and reranking candidate count affect cost efficiency and latency requirements.





