OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests
The Next Web reported that OpenAI has built GPT-Red, an internal automated red-team model for prompt-injection attacks, but is keeping the attacker private. The report cited attack success rates above 90% against an older GPT-5 and below 23% against GPT-5.6, while noting that human testers still catch cases GPT-Red misses.

OpenAI's internal automated red-team model, GPT-Red, is built to attack other AI systems, and The Next Web indicated that it is not being released publicly.
The tool focuses on prompt injection, where hidden instructions in files, webpages or messages can push an AI agent into actions its operator did not intend.
The organization characterizes GPT-Red as a safety system rather than a public product.
Its training is designed to uncover ways to hijack or sabotage AI systems so developers can patch weaknesses before release, The Next Web wrote.
GPT-Red Targets Prompt-Injection Attacks
Through self-play against defender models, GPT-Red learns to attack, according to the outlet.
The attacker receives rewards when an exploit succeeds, and the defenders receive rewards for blocking the attempt.
For this safety work, OpenAI used some of its largest compute runs, the outlet noted.
The account did not put a dollar cost or hardware count on the training run.
Researchers also identified a prompt-injection method that OpenAI described as a fake chain of thought.
Chris Choquette-Choo told MIT Technology Review that the attack plants a false note in a model's private working memory, making the target treat an untrue statement as already verified.
Vendy Test Gave GPT-Red A Physical Target
One test moved beyond text-only examples.
In that trial, GPT-Red attacked Vendy, an AI agent built by Andon Labs to run a real vending machine in OpenAI's office, The Next Web wrote.
OpenAI said the attack changed prices, pushed one expensive item down to the 50-cent minimum and cancelled a customer order.
OpenAI disclosed the flaws after the test.
The result set separated older GPT-5, GPT-5.6 and human testers by large margins.
The Next Web wrote that more than 90% of GPT-Red's strongest attacks worked against an older GPT-5, while fewer than 23% succeeded against GPT-5.6.
The same account added that a rerun of a 2025 test had GPT-Red cracking 84% of scenarios compared with 13% for human red-teamers.
Human Testers Still Catch Missed Cases
Against GPT-Red, OpenAI trained GPT-5.6 and described the newer model as its most robust system against prompt injection, according to the outlet.
It is using the attacker to harden later models rather than distributing it as a tool for outside researchers.
Extended back-and-forth attacks and attempts that hide instructions in images remain weaker areas for GPT-Red, The Next Web noted.
Jessica Ji, an AI security analyst at Georgetown's CSET, told the outlet that human expertise will still be very important.
OpenAI has not released GPT-Red.
TNW did not report external researcher access terms for GPT-Red.




















