# Microsoft — Beyond violation rates: Are your AI evaluations measuring the right things?

- Company: Microsoft (microsoft.com)
- Announced: 2026-10-05T14:54:13+00:00
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://commandline.microsoft.com/structured-ai-evaluation-design-assert-vs-petri-bloom/
- Record: https://forck.live/items/19122-beyond-violation-rates-are-your-ai-evaluations-measuring-the-right-things
- Subject: Foundry / Responsible AI

Microsoft's research compares its ASSERT framework with Petri Bloom for generating evaluation suites for generative AI safety. The study examines coverage, effectiveness, and robustness across 16 cybersecurity risks, finding that evaluation quality depends on more than pass/fail rates. ASSERT organizes tests around an explicit taxonomy to make coverage inspectable and connect failures to specific behaviors.

## Evidence

Verbatim from https://commandline.microsoft.com/structured-ai-evaluation-design-assert-vs-petri-bloom/:

> Evaluation quality depends not only on how often tests pass or fail, but also on what the suite covers, how efficiently it exposes failures, and whether its conclusions survive changes to judging and test construction.

---

Record: https://forck.live/items/19122-beyond-violation-rates-are-your-ai-evaluations-measuring-the-right-things
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
