# Amazon — Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

- Company: Amazon (amazon.com)
- Announced: 2026-08-26T19:13:35+00:00
- Category: capability-change
- Subject: Bedrock / Nova
- Source: https://aws.amazon.com/blogs/machine-learning/evaluate-any-agent-framework-with-amazon-bedrock-agentcore-evaluations/
- Record: https://forck.live/items/4475-evaluate-any-agent-framework-with-amazon-bedrock-agentcore-evaluations

Amazon Bedrock AgentCore Evaluations is a new evaluation service that uses OpenTelemetry to score any agent framework (LangGraph, LlamaIndex, OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands Agents) running on Amazon Bedrock AgentCore runtime, without requiring framework-specific integration. It reads three span roles (invoke agent, inference, execute tool) to reconstruct sessions and apply evaluators such as GoalSuccessRate, Correctness, Helpfulness, and custom LLM-as-a-judge. The service bridges both OpenTelemetry GenAI conventions and OpenInference specifications.

## Evidence

Verbatim from https://aws.amazon.com/blogs/machine-learning/evaluate-any-agent-framework-with-amazon-bedrock-agentcore-evaluations/:

> Amazon Bedrock AgentCore evaluations solves this fragmentation by decoupling evaluation from the framework choice.

---

Record: https://forck.live/items/4475-evaluate-any-agent-framework-with-amazon-bedrock-agentcore-evaluations
Catalogue: https://forck.live/llms.txt
Feed: https://forck.live/feed.md
