From the source
Cua announced an integration with HUD, an evaluation platform for computer-use agents, allowing users to benchmark any GUI-capable agent on real computer-use tasks.
The integration supports one-line evaluations on benchmarks like OSWorld and SheetBench for models from OpenAI, Anthropic, Hugging Face, and composed GUI models, with live traces available at app.hud.ai.





