# Amazon — Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

- Company: Amazon (amazon.com)
- Announced: 2026-08-27T16:05:10+00:00
- Category: capability-change
- Subject: Bedrock / Nova
- Models affected: Parakeet TDT 0.6B V2, Parakeet TDT 0.6B V2
- Source: https://aws.amazon.com/blogs/machine-learning/reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2/
- Record: https://forck.live/items/4517-reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2

Amazon Web Services, NVIDIA, and Heidi Health announce a solution that uses NVIDIA CUDA Multi-Process Service and NVIDIA Triton Inference Server on Amazon EC2 GPU instances to reduce GPU infrastructure requirements for ASR inference by 75 percent, from 16 instances to 4, while maintaining sub-second latency at 92.1 requests per second per GPU, enabling Heidi Health to process over 2.4 million clinical consultations per week.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2/:

> In this post, we focus on what comes after fine-tuning: serving that model efficiently. We demonstrate how NVIDIA CUDA Multi-Process Service (MPS), combined with NVIDIA Triton Inference Server™ on Amazon EC2 GPU instances, reduces GPU infrastructure requirements by 75 percent (from 16 instances to 4).

---

Record: https://forck.live/items/4517-reduce-asr-inference-costs-by-75-with-nvidia-mps-on-amazon-ec2
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
