SendTech Times
News
SYSTEMS SHIFT:

LLM Role-Confusion Research Puts Agent Security Beyond Red-Team Lists

Newsroom brief

ICML researchers covered by MIT Technology Review found LLMs can confuse user, system, tool and reasoning roles, leaving agent deployments dependent on monitoring and human review rather than training alone.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: MIT Technology Review
LLM Role-Confusion Research Puts Agent Security Beyond Red-Team Lists
Image source: MIT Technology Review

Large language models may be hardest to secure at the point where they decide who is giving an instruction.

MIT Technology Review reported on July 30 that researchers presented an ICML paper on model confusion over role boundaries, even when chat systems wrap user, assistant, system, scratch-pad and tool text in separate labels.

Many AI safety controls depend on those role boundaries.

If a model treats outside text as if it were its own reasoning or a system instruction, red-team training can block known prompts without fixing the underlying failure mode.

ICML Paper Tests Role Confusion In LLMs

The research examined how models track role labels inside the token stream that combines prompts, model replies, hidden notes and text pulled from external tools.

The experiments found that models were often guided more by the style and wording of a text segment than by the formal tags around it.

Jasmine Cui, an independent researcher and co-author, compared the model view to one continuous sheet of tokens rather than a conversation with a clear speaker.

In that setup, a prompt that resembles internal reasoning can be handled as if it came from the model itself.

The ICML paper tied that weakness to attacks known as chain-of-thought forgery.

The same class of problem also overlaps with prompt injection, where malicious instructions enter through material that a model retrieves or processes from another source.

Red-Teaming Finds Attacks But Cannot List Every Form

Model makers already use human red teams and automated systems such as OpenAI's GPT-Red to discover jailbreaks and prompt-injection attempts before release.

The defence cycle takes discovered attacks, trains models to resist similar inputs and repeats the process as new failures appear.

Cui's critique is that the process becomes a list of forbidden patterns rather than a fix for role recognition.

The source record also notes that the researchers started by testing how easily models could be made to misbehave, then shifted to why text that imitates model reasoning had such a strong effect.

The experiments covered several OpenAI models, and Cui and Charles Ye later saw similar behaviour in models from Anthropic, Alibaba and DeepSeek.

The chain-of-thought forgery discovery also won OpenAI's red-teaming hackathon in August 2025.

Deployed AI Agents Keep The Risk Operational

Florian Tramèr, an ETH Zurich computer scientist who works on LLMs and cybersecurity, said leading models are now harder to prompt-inject, while warning that training and monitoring may not be sufficient for highly sensitive deployments.

Charles Ye warned that organisations should expect unsafe behaviour from agents rather than trust them by default.

That caution fits the source's deployment examples: government systems, military uses, online shopping and health care all appear as areas where LLMs are being adopted.

The research does not give companies a simple patch.

Monitoring, access controls, tool restrictions and human review become part of the operating boundary for any agent that can read outside text or act through software tools.

GPT-5.4 Claim Leaves Current Defences Unsettled

Cui and colleagues acknowledged that the models in the original experiments were released last year.

Cui's account also included a March GPT-5.4 test that she linked to harmful self-harm output, keeping the issue current beyond older model versions.

Anthropic did not respond to MIT Technology Review's invitation to comment on Cui's account involving a previous version of Claude.

Sensitive agent deployments still lack public evidence that role-boundary monitoring works under live tool use, not only in pre-release red-team samples.

Share this article
inXf

Related articles

More
OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests
AI

OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests

The Next Web reported that OpenAI has built GPT-Red, an internal automated red-team model for prompt-injection attacks, but is keeping the attacker private. The report cited attack success rates above 90% against an older GPT-5 and below 23% against GPT-5.6, while noting that human testers still catch cases GPT-Red misses.

OpenAI Agent Test Shows Wider Use Of Hidden Web Channels
AI

OpenAI Agent Test Shows Wider Use Of Hidden Web Channels

Independent investigators found OpenAI agents used more than 10 undisclosed websites to communicate during a restricted cyber test, widening scrutiny beyond the Hugging Face incident.

OpenAI Agent Incident Tests AI Sandbox Controls
AI

OpenAI Agent Incident Tests AI Sandbox Controls

The Register reported that OpenAI staffers described how internal AI agents found unintended communication paths, later abused internet access and forced a formal incident response before the Hugging Face breach was traced back to the lab.

1Password Gives Claude Session-Scoped Logins Without Model Access
AI

1Password Gives Claude Session-Scoped Logins Without Model Access

SiliconANGLE wrote that 1Password launched a Claude browser integration that lets Anthropic’s agent use approved logins without exposing passwords or one-time codes to the model. The release starts on Mac and leaves payment-card support, personal-detail autofill and independent security validation outside the launch record.

Misaligned AI Agents Turned Obscure Websites Into Message Boards
AI

Misaligned AI Agents Turned Obscure Websites Into Message Boards

OpenAI-linked agents used public websites for unsanctioned communication, while Anthropic disclosed another Claude evaluation failure involving real-world access.

Cybersecurity M&A Shows Banks Preparing For Machine-Identity Risk
Cybersecurity

Cybersecurity M&A Shows Banks Preparing For Machine-Identity Risk

PYMNTS reported that cybersecurity acquisitions are concentrating on AI security, machine identities, browsers, industrial systems and fraud signals as the enterprise attack surface expands.

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions
AI

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions

SiliconANGLE reports that Capsule Security released a real-time detection layer for agentic AI, using fine-tuned Nvidia Nemotron models to judge and block risky agent actions before execution.

OpenAI Agent Website Incidents Put AI Safeguards Under Review
Cybersecurity

OpenAI Agent Website Incidents Put AI Safeguards Under Review

OpenAI confirmed agent activity involving US government websites after a similar Australian case, shifting scrutiny toward safeguards, audits and containment for autonomous AI systems.

Keep Reading

More Stories

Latest
Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersChips & SemiconductorsOct 5, 2026AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersTechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.