AI Agents in Offensive Security: What the Tests Actually Revealed
In a development that has sent ripples through the cybersecurity and AI policy communities, AI agents developed by OpenAI and Anthropic were used in controlled red-team exercises where they targeted real people and live systems — not just sandboxed simulations. The AI agents cybersecurity testing scenarios, reported by BleepingComputer, represent a significant leap in how frontier AI labs are stress-testing their models' capabilities, and they raise profound questions about the governance, containment, and regulation of autonomous AI systems.
The exercises were structured as offensive security tests — the kind commonly run by penetration testers and red teams — but with a crucial difference: the agents were AI-driven, capable of autonomous decision-making, and in some cases, reportedly engaged with real individuals and operational infrastructure. Security researchers and privacy professionals are now grappling with the implications, particularly as the European Union's AI Act continues to take shape and regulators worldwide seek frameworks for governing AI systems that can act independently in high-stakes environments.
For IT decision-makers, developers, and privacy professionals, these tests are not merely academic. They illustrate a capability threshold that changes the threat landscape — and the compliance landscape — fundamentally.
What Happened During the AI Red-Team Exercises?

Red-teaming is a long-established practice in cybersecurity, where security professionals simulate adversarial attacks to identify vulnerabilities before real attackers do. What makes the OpenAI and Anthropic tests significant is the substitution — or augmentation — of human operators with autonomous AI agents capable of planning, executing, and adapting multi-step attacks.
According to the original reporting from BleepingComputer, both companies ran scenarios in which their AI agents interacted with real people and real systems as part of controlled cyber evaluations. This is a notable departure from purely synthetic testing environments. The agents demonstrated an ability to conduct social engineering, gather intelligence, and probe system defenses — capabilities that, outside a controlled context, would constitute serious offensive cyber operations.
These findings align with a growing body of research. A study published on arXiv by researchers at the University of Illinois Urbana-Champaign found that GPT-4-class agents could autonomously exploit real-world vulnerabilities, including one-day CVEs, with a success rate that alarmed the security community. The researchers noted that the agents required no prior knowledge of the specific vulnerabilities to succeed — they could reason their way to exploitation.
"The transition from AI-assisted to AI-autonomous offense is not a future concern — it is a present capability that security teams must account for today."
— Senior red-team researcher, speaking at a recent AI security symposiumAnthropic, which has positioned its Claude models around "Constitutional AI" and safety-first design principles, has been particularly open about red-teaming as a core part of its model evaluation process. The company's research into "model organisms of misalignment" has explored how AI systems might behave deceptively or pursue unintended goals under certain conditions. OpenAI similarly maintains a Safety Systems team and has published documentation on how it evaluates dangerous capabilities ahead of model releases.
The Detection Gap: Why 54% Success Rates Should Alarm Every Security Team
A critical data point from the security research community puts the AI agent testing story into sharp relief. According to Picus Security's research on breach and attack simulation, security teams successfully log only 54% of attacks that reach their environment — and generate meaningful alerts on just 14% of threats. This means that the vast majority of sophisticated intrusions move through corporate and government networks entirely unseen by SIEM and EDR systems.
Now layer on top of that the capability demonstrated in these AI agent tests: an autonomous system that can adapt its tactics in real time, learn from failed attempts, conduct social engineering at scale, and operate across multiple attack vectors simultaneously. The detection gap does not merely persist — it widens dramatically when the adversary is an AI agent rather than a human attacker constrained by time, fatigue, and cognitive bandwidth.
Breach and attack simulation (BAS) platforms have emerged as one response to this challenge, continuously testing SIEM detection rules and EDR policies against simulated attack chains. But as Wired has reported in its coverage of AI-enabled offensive capabilities, BAS platforms themselves will need to incorporate AI-adversary modeling to remain relevant. A red-team exercise conducted with a static script is a fundamentally different threat model than one conducted by an adaptive AI agent.
AI Regulation and Digital Sovereignty: Europe's High-Stakes Response

For European policymakers and the privacy professionals who operate within the EU's regulatory ecosystem, these AI agent tests arrive at a particularly consequential moment. The EU AI Act — the world's first comprehensive legal framework for artificial intelligence — classifies systems that could threaten critical infrastructure or be used for offensive cyber operations as high-risk. Under its provisions, such systems require mandatory conformity assessments, transparency obligations, and human oversight mechanisms before deployment.
But the tests conducted by OpenAI and Anthropic expose a regulatory blind spot: what is the status of an AI agent that is used in a controlled red-team exercise, operates on real infrastructure, and engages with real individuals — but under the supervision of the AI company itself? The answer is far from clear, and legal scholars across Europe are already debating whether existing provisions adequately capture the risk profile of frontier AI agents used in offensive security research.
The GDPR adds another layer of complexity. If an AI agent interacts with real people during a security test — gathering personal data, crafting targeted communications, or probing systems that contain personal information — GDPR obligations may be triggered regardless of the research context. Data minimisation, purpose limitation, and lawful basis requirements do not come with research exemptions broad enough to cover all conceivable red-team scenarios, particularly when non-consenting third parties are involved.
| Regulatory Framework | Relevance to AI Agent Cyber Testing | Key Requirement |
|---|---|---|
| EU AI Act | High-risk classification for offensive cyber AI | Mandatory conformity assessment, human oversight |
| GDPR | Data processing during tests involving real people | Lawful basis, data minimisation, purpose limitation |
| NIS2 Directive | Cybersecurity obligations for critical infrastructure operators | Incident reporting, risk management measures |
| Cyber Resilience Act | Security requirements for products with digital elements | Security by design, vulnerability disclosure |
| DORA (Financial Sector) | Threat-led penetration testing requirements | Controlled testing of live systems under strict oversight |
The European approach to digital sovereignty — the principle that EU citizens and institutions should maintain control over their data and digital infrastructure — is directly challenged by AI systems that can autonomously probe, map, and engage with that infrastructure. Open-source alternatives and self-hosted AI deployments are gaining traction in the European enterprise and public sector market partly in response to precisely these concerns.
What Security Teams and IT Decision-Makers Must Do Now
For CISOs, IT managers, and security architects, the lessons from these AI agent tests are operational as well as strategic. The emergence of autonomous AI as a credible offensive capability demands a reassessment of threat models, detection architectures, and incident response playbooks.
First, the assumption that human-speed attacks define the threat ceiling must be abandoned. AI agents can operate around the clock, across multiple attack vectors simultaneously, and adapt their approach based on real-time feedback. A social engineering campaign that would take a human red-teamer days to execute can be compressed into hours or minutes by an autonomous agent. Detection rules tuned to human behavioral patterns may simply not flag AI-driven activity that operates at different cadences and with different characteristics.
Second, the detection gap highlighted by Picus Security's research — 54% logging, 14% alerting — represents an existential problem in an AI-adversary threat environment. Security teams should be running regular breach and attack simulation exercises specifically designed to test whether their SIEM and EDR infrastructure would detect AI-agent-style attack chains. MITRE ATT&CK framework mappings need to be reviewed for AI-specific tactics, techniques, and procedures (TTPs) that may not yet be well-represented.
Third, organisations operating under GDPR or processing EU citizens' data must begin asking hard questions about their own use of AI security tools. The same AI agent capabilities that make offensive testing powerful also power defensive tools — and those defensive tools may themselves create data protection obligations depending on how they collect, process, and retain information about individuals or systems.