From the source
Microsoft's research compares its ASSERT framework with Petri Bloom for generating evaluation suites for generative AI safety.
The study examines coverage, effectiveness, and robustness across 16 cybersecurity risks, finding that evaluation quality depends on more than pass/fail rates.
ASSERT organizes tests around an explicit taxonomy to make coverage inspectable and connect failures to specific behaviors.





