From the source
Cua released a research Metal capability shim that intercepts GPU capability queries inside macOS VMs, allowing llama.cpp to select newer Metal kernels.
On an M1 Ultra, TinyLlama 1.1B prompt processing improved 11.08× and token generation 16.36× versus a stock VM.
The shim is released under a permissive license for reproduction and further testing.






