From the source
To deploy large language models like BLOOMZ on Habana Gaudi2 accelerators using Optimum Habana, and presents benchmarks showing that Gaudi2 achieves lower inference latency than Nvidia A100 80GB GPUs for BLOOMZ models.
From the source
From the source

To deploy large language models like BLOOMZ on Habana Gaudi2 accelerators using Optimum Habana, and presents benchmarks showing that Gaudi2 achieves lower inference latency than Nvidia A100 80GB GPUs for BLOOMZ models.
From the source
Gaudi2 is 2.89x faster than A100 for BLOOMZ-7B!
huggingface.co