AWS Workflow Lets GitHub Actions Block Failed AI Agent Checks
AWS has published a reference workflow for testing AI agents in GitHub Actions, using AgentCore Evaluations to score traces, block risky merges and expose runtime and judge-model tradeoffs.

DeveloperTech reported that AWS has released a reference setup that lets GitHub Actions test AI agents before code is merged, moving agent evaluation into the same CI/CD lane as other software checks.
The workflow is designed to turn a pull request into an agent test run.
When a pull request touches the agent, its infrastructure, an MCP server or the test scripts, the job creates a development stack, runs preset prompts, gathers telemetry and sends scores back into CI.
Those scores can become a merge gate.
Branch protections or rulesets can make the job mandatory, turning a failed agent score into a blocked merge rather than a warning outside the review flow.
AWS built the sample around Amazon Bedrock AgentCore Evaluations and a Strands-based agent running on AgentCore Runtime.
In the sample stack, the agent reaches tools through a Model Context Protocol server, and AWS Cloud Development Kit creates the runtime pieces, Cognito setup and IAM roles.
The September 8 guide uses four built-in checks: GoalSuccessRate, Correctness, ToolSelectionAccuracy and ToolParameterAccuracy.
Together they examine whether the agent finished the requested task, produced a suitable answer, chose the right tool and derived tool parameters from the conversation context.
The reference threshold is 0.8.
In AWS's deliberate regression example, a changed system prompt made GoalSuccessRate, Correctness and ToolParameterAccuracy fall below the line, while ToolSelectionAccuracy still passed.
Restoring the original prompt and rerunning the workflow brought all four evaluations back above the threshold.
The setup also points to a wider testing model for agent teams.
AgentCore dataset evaluation can run collections of scenarios with expected responses, natural-language assertions and expected trajectories for tool calls.
Correctness can compare output with an expected response, GoalSuccessRate can test session-level assertions, and trajectory evaluators can compare the actual execution path with the intended sequence.
That matters because AWS has warned that large language models can return different outputs, tool choices and paths across repeated runs.
Reusable datasets give teams a way to rerun expected requests and behaviours instead of judging only a final answer from one session.
AgentCore can score more than the final response.
Tests may target a full interaction, one recorded trace or selected tool-call spans; its modes cover controlled CI runs, sampled production behaviour and asynchronous scoring of multiple sessions.
The reference pipeline relies on OpenTelemetry data generated during agent sessions.
AWS's starter toolkit retrieves the relevant traces from Amazon CloudWatch before submitting them for scoring, and a second documented option lets teams evaluate stored staging traces in CI without invoking a live AgentCore runtime during the pull-request check.
The implementation carries operational limits as well as test coverage.
GitHub Actions uses OIDC federation in the sample to assume an AWS IAM role without storing long-lived credentials, and Cognito supplies a machine-to-machine token for OAuth-protected MCP tools.
AWS also notes that the sample gives the CI pipeline access to all tools rather than testing role-specific access controls.
Runtime and cost are part of the tradeoff.
AWS puts the full sample run at roughly 10 minutes, with deployment, startup, trace availability and scoring all included in the CI job.
Trace propagation took between 30 and 90 seconds in AWS's tests, with retries every 30 seconds for as long as 10 minutes.
LLM-based evaluators add another variable.
Four evaluators against five prompts create 20 judge-model calls for each pull request in the example, and AWS warns that LLM-as-a-judge scores can vary across repeated evaluations.
Trajectory checks avoid those model calls by comparing recorded execution paths with predefined ground truth, while Lambda-based custom evaluators can add deterministic checks for application-specific requirements.




















