From the source
Cua evaluated Gemini 3.5 Flash with its native Computer Use API on the Cua-Bench KiCad EDA suite of 25 real electrical-engineering tasks at a 200-step budget.
The model achieved the highest mean reward (0.267) among tested frontier models, solving 5 tasks fully and 3 partially, while GPT-5.5 solved 6 tasks outright but earned no partial credit.
The evaluation highlighted strengths in pixel-accurate grounding on zoomed-in targets and analog-design reasoning, with main losses from design-from-scratch tasks timing out and occasional hallucination of screen state after the second screenshot.






