From the source
Hugging Face publishes a blog post benchmarking BERT-like model inference on modern CPUs, covering hardware optimizations, core scaling, and batch size scaling using PyTorch, TensorFlow, TorchScript, XLA, and ONNX Runtime on an AWS c5.metal instance with an Intel Xeon Platinum 8275 CPU.




