# Amazon — Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

- Company: Amazon (amazon.com)
- Announced: 2026-10-05T15:50:01+00:00
- Category: capability-change
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://aws.amazon.com/blogs/machine-learning/evaluating-multi-agent-systems-for-explainability-and-helpfulness-with-amazon-bedrock-agentcore/
- Record: https://forck.live/items/16056-evaluating-multi-agent-systems-for-explainability-and-helpfulness-with-amazon
- Subject: Bedrock / Nova

Amazon Bedrock AgentCore Evaluations, a fully managed capability of Amazon Bedrock AgentCore, enables teams to assess agent performance across development and production by measuring accuracy, task success, and behavior across multiple quality dimensions. It supports both built-in evaluators for common dimensions like helpfulness and task success, and custom evaluators for domain-specific validation such as explainability and constraint satisfaction. The solution also integrates with Amazon Bedrock Guardrails for safety controls during execution.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/evaluating-multi-agent-systems-for-explainability-and-helpfulness-with-amazon-bedrock-agentcore/:

> Amazon Bedrock AgentCore Evaluations , a capability of Amazon Bedrock AgentCore, is designed to address this challenge as a fully managed capability for assessing agent performance across development and production, so teams can measure accuracy, task success, and behavior across multiple quality dimensions.

---

Record: https://forck.live/items/16056-evaluating-multi-agent-systems-for-explainability-and-helpfulness-with-amazon
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
