SendTech Times
News
SYSTEMS SHIFT:

AWS Workflow Lets GitHub Actions Block Failed AI Agent Checks

Newsroom brief

AWS has published a reference workflow for testing AI agents in GitHub Actions, using AgentCore Evaluations to score traces, block risky merges and expose runtime and judge-model tradeoffs.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: DeveloperTech
AWS Workflow Lets GitHub Actions Block Failed AI Agent Checks
Image source: DeveloperTech

DeveloperTech reported that AWS has released a reference setup that lets GitHub Actions test AI agents before code is merged, moving agent evaluation into the same CI/CD lane as other software checks.

The workflow is designed to turn a pull request into an agent test run.

When a pull request touches the agent, its infrastructure, an MCP server or the test scripts, the job creates a development stack, runs preset prompts, gathers telemetry and sends scores back into CI.

Those scores can become a merge gate.

Branch protections or rulesets can make the job mandatory, turning a failed agent score into a blocked merge rather than a warning outside the review flow.

AWS built the sample around Amazon Bedrock AgentCore Evaluations and a Strands-based agent running on AgentCore Runtime.

In the sample stack, the agent reaches tools through a Model Context Protocol server, and AWS Cloud Development Kit creates the runtime pieces, Cognito setup and IAM roles.

The September 8 guide uses four built-in checks: GoalSuccessRate, Correctness, ToolSelectionAccuracy and ToolParameterAccuracy.

Together they examine whether the agent finished the requested task, produced a suitable answer, chose the right tool and derived tool parameters from the conversation context.

The reference threshold is 0.8.

In AWS's deliberate regression example, a changed system prompt made GoalSuccessRate, Correctness and ToolParameterAccuracy fall below the line, while ToolSelectionAccuracy still passed.

Restoring the original prompt and rerunning the workflow brought all four evaluations back above the threshold.

The setup also points to a wider testing model for agent teams.

AgentCore dataset evaluation can run collections of scenarios with expected responses, natural-language assertions and expected trajectories for tool calls.

Correctness can compare output with an expected response, GoalSuccessRate can test session-level assertions, and trajectory evaluators can compare the actual execution path with the intended sequence.

That matters because AWS has warned that large language models can return different outputs, tool choices and paths across repeated runs.

Reusable datasets give teams a way to rerun expected requests and behaviours instead of judging only a final answer from one session.

AgentCore can score more than the final response.

Tests may target a full interaction, one recorded trace or selected tool-call spans; its modes cover controlled CI runs, sampled production behaviour and asynchronous scoring of multiple sessions.

The reference pipeline relies on OpenTelemetry data generated during agent sessions.

AWS's starter toolkit retrieves the relevant traces from Amazon CloudWatch before submitting them for scoring, and a second documented option lets teams evaluate stored staging traces in CI without invoking a live AgentCore runtime during the pull-request check.

The implementation carries operational limits as well as test coverage.

GitHub Actions uses OIDC federation in the sample to assume an AWS IAM role without storing long-lived credentials, and Cognito supplies a machine-to-machine token for OAuth-protected MCP tools.

AWS also notes that the sample gives the CI pipeline access to all tools rather than testing role-specific access controls.

Runtime and cost are part of the tradeoff.

AWS puts the full sample run at roughly 10 minutes, with deployment, startup, trace availability and scoring all included in the CI job.

Trace propagation took between 30 and 90 seconds in AWS's tests, with retries every 30 seconds for as long as 10 minutes.

LLM-based evaluators add another variable.

Four evaluators against five prompts create 20 judge-model calls for each pull request in the example, and AWS warns that LLM-as-a-judge scores can vary across repeated evaluations.

Trajectory checks avoid those model calls by comparing recorded execution paths with predefined ground truth, while Lambda-based custom evaluators can add deterministic checks for application-specific requirements.

Share this article
inXf

Related articles

More
AWS Adds Persistent Runtime Instances For Production AI Agents
Cloud & Data Centers

AWS Adds Persistent Runtime Instances For Production AI Agents

AWS announced runtime instances for Amazon Bedrock AgentCore Runtime, adding managed infrastructure for multi-agent workflows, shared sessions lasting up to 14 days and GPU-supported production agent deployments.

Microsoft Deploys 25 AI Agents Across Supply-Chain Workflows
AI

Microsoft Deploys 25 AI Agents Across Supply-Chain Workflows

Microsoft has put more than 25 AI agents into supply-chain operations, using them to support logistics planning, freight auditing, spare-parts availability and invoice checks.

Shopify Reworks Theme Code For AI Agents And Readability
AI

Shopify Reworks Theme Code For AI Agents And Readability

Shopify is preparing a cleaner storefront theme that shifts from JSON-heavy configuration toward mostly HTML and Liquid, as AI-assisted theme editing makes readable code and explicit contracts more important.

GitHub And Google Back ARD As AI Agents Search For Tools
AI

GitHub And Google Back ARD As AI Agents Search For Tools

GitHub, Google, Microsoft and other companies are backing Agentic Resource Discovery, a specification meant to help AI agents find, verify and connect to tools, skills, MCP servers and other resources without hard-coded integrations.

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions
AI

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions

SiliconANGLE reports that Capsule Security released a real-time detection layer for agentic AI, using fine-tuned Nvidia Nemotron models to judge and block risky agent actions before execution.

Novo Nordisk And AWS Build Agentic AI Workflows For Drug Discovery
AI

Novo Nordisk And AWS Build Agentic AI Workflows For Drug Discovery

AI News reported that Novo Nordisk named AWS its preferred cloud provider and strategic AI partner, adding a London co-innovation hub and agentic AI tools for drug-target, therapy-design and research workflows.

AI Agent Rollouts Require Testing Before Live Customers
AI

AI Agent Rollouts Require Testing Before Live Customers

No Jitter reported that enterprise AI agents need guardrails, simulations, answer checks and visibility before customer-facing deployment, as vendors add tools to catch regressions and rollback failures.

Amazon Blocks Meta Muse From Shopping As Agent Controls Tighten
AI

Amazon Blocks Meta Muse From Shopping As Agent Controls Tighten

Amazon blocked Meta’s Muse AI agent from shopping on its platform, raising retailer control, privacy and credential questions around consumer buying agents.

Keep Reading

More Stories

Latest
Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersChips & SemiconductorsOct 5, 2026AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersTechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.