# Cua — Evaluating Gemini 3.5 Flash on Computer-Use with Cua-Bench

- Company: Cua (cua.ai)
- Announced: 2026-06-24
- Category: not stated
- Coverage: not counted
- Announcement: no
- Group: routine
- Source: https://cua.ai/blog/evaluating-gemini-3.5-flash-on-computer-use
- Record: https://forck.live/items/16819-evaluating-gemini-3-5-flash-on-computer-use-with-cua-bench
- Subject: Cua Driver / Lume / Cua Bench / Cua Fleets
- Models affected: gemini-3.5-flash, GPT-5.5, Claude Sonnet 4.5, Claude Haiku 4.5, Claude Opus 4.8, Gemini 3.1 Pro, Gemini 3 Flash

Cua evaluated Gemini 3.5 Flash with its native Computer Use API on the Cua-Bench KiCad EDA suite of 25 real electrical-engineering tasks at a 200-step budget. The model achieved the highest mean reward (0.267) among tested frontier models, solving 5 tasks fully and 3 partially, while GPT-5.5 solved 6 tasks outright but earned no partial credit. The evaluation highlighted strengths in pixel-accurate grounding on zoomed-in targets and analog-design reasoning, with main losses from design-from-scratch tasks timing out and occasional hallucination of screen state after the second screenshot.

## Evidence

Verbatim from https://cua.ai/blog/evaluating-gemini-3.5-flash-on-computer-use:

> Gemini 3.5 Flash posted the highest mean reward of any frontier model we tested: 0.267 , with 5 of 25 tasks solved fully and 3 more partially.

---

Record: https://forck.live/items/16819-evaluating-gemini-3-5-flash-on-computer-use-with-cua-bench
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
