OpenAI Keeps GPT-Red Attack Model Private After Prompt-Injection Tests
The Next Web reported that OpenAI has built GPT-Red, an internal automated red-team model for prompt-injection attacks, but is keeping the attacker private. The report cited attack success rates above 90% against an older GPT-5 and below 23% against GPT-5.6, while noting that human testers still catch cases GPT-Red misses.

OpenAI's internal automated red-team model, GPT-Red, is built to attack other AI systems and is not being released publicly, The Next Web reported.
The work focuses on prompt injection, where hidden instructions in files, webpages or messages can push an AI agent into actions its operator did not intend.
The tension is clear: the same capability that helps OpenAI find agent failures could also become a tool for abusing other systems if distributed too widely.
GPT-Red is therefore being positioned as internal safety infrastructure, not as an outside developer product.
GPT-Red Targets Prompt-Injection Attacks
Through self-play against defender models, GPT-Red learns to attack.
The attacker receives rewards when an exploit succeeds, and the defenders receive rewards for blocking the attempt.
OpenAI used some of its largest compute runs for the safety work, but TNW did not put a dollar cost or hardware count on the training run.
The absence of those figures matters because automated red-teaming is becoming a compute problem as well as a security problem: stronger attack models can test more scenarios, but they may also concentrate safety tooling among companies with the infrastructure to train them.
Researchers also identified a prompt-injection method that OpenAI described as a fake chain of thought.
Chris Choquette-Choo told MIT Technology Review that the attack plants a false note in a model's private working memory, making the target treat an untrue statement as already verified.
Vendy Test Gave GPT-Red A Physical Target
One test moved beyond text-only examples.
GPT-Red attacked Vendy, an AI agent built by Andon Labs to run a real vending machine in OpenAI's office.
OpenAI said the attack changed prices, pushed one expensive item down to the 50-cent minimum and cancelled a customer order.
That makes the test more than a lab puzzle: it shows how prompt injection can reach pricing, order handling and other operational controls when AI agents are connected to real systems.
The result set separated older GPT-5, GPT-5.6 and human testers by large margins.
TNW wrote that more than 90% of GPT-Red's strongest attacks worked against an older GPT-5, while fewer than 23% succeeded against GPT-5.6.
The same account added that a rerun of a 2025 test had GPT-Red cracking 84% of scenarios compared with 13% for human red-teamers.
Human Testers Still Catch Missed Cases
OpenAI trained GPT-5.6 against GPT-Red and described the newer model as its most robust system against prompt injection.
The company is using the attacker to harden later models rather than distributing it as a tool for outside researchers.
Extended back-and-forth attacks and attempts that hide instructions in images remain weaker areas for GPT-Red.
Jessica Ji, an AI security analyst at Georgetown's CSET, told TNW that human expertise will still be very important.
OpenAI has not released GPT-Red, and TNW did not report external researcher access terms.
For companies deploying agents, the immediate lesson is narrower: prompt-injection testing has to cover connected workflows, not only chat responses, because a successful attack can move from hidden text into prices, orders and other live actions.




















