# Cua — Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with llama.cpp

- Company: Cua (cua.ai)
- Announced: 2026-08-11
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://cua.ai/blog/gpu-passthrough-macos-vms
- Record: https://forck.live/items/16816-apple-silicon-and-macos-vms-11-16-faster-llm-inference-with-llama-cpp
- Subject: Cua Driver / Lume / Cua Bench / Cua Fleets
- License: same permissive license as Lume and Cua

Cua released a research Metal capability shim that intercepts GPU capability queries inside macOS VMs, allowing llama.cpp to select newer Metal kernels. On an M1 Ultra, TinyLlama 1.1B prompt processing improved 11.08× and token generation 16.36× versus a stock VM. The shim is released under a permissive license for reproduction and further testing.

## Evidence

Verbatim from https://cua.ai/blog/gpu-passthrough-macos-vms:

> On an M1 Ultra, TinyLlama 1.1B running through llama.cpp processed prompts 11.08× faster and generated tokens 16.36× faster than the same workload in the same stock VM.

---

Record: https://forck.live/items/16816-apple-silicon-and-macos-vms-11-16-faster-llm-inference-with-llama-cpp
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
