From the source
Cua compiled a curated list of 45 computer-use agent papers presented at NeurIPS 2025, covering benchmarks, safety, grounding, visual reasoning, and agent architectures.
The post highlights that the benchmark landscape is maturing with evaluations across macOS, professional tools, and real-world websites, and that safety is becoming a first-class concern with multiple papers documenting agent failures under adversarial inputs, privacy requirements, or misuse scenarios.






