From the source
Performance benchmarks of Hugging Face Infinity on Intel Xeon CPUs, showing up to 34% better latency/throughput compared to previous generation and up to 800% better than vanilla transformers.
It also notes that Infinity is no longer offered as a commercial inference solution as of December 2022.






