Autonomous AI agents now plan, adapt, and chain exploits in penetration tests at scale. From depthfirst integrations to Fleuret AI's continuous defense, vendors race to close the gap between attacker speed and defender response. Yet governance, proof of impact, and human oversight remain essential. The technology is real and advancing fast.
Security teams have long accepted a painful trade-off. Manual penetration tests deliver high-quality findings but happen too rarely and cover too little ground. Automated scanners run often yet miss complex, chained vulnerabilities that real attackers exploit. Now a new approach promises to split that difference.
Agentic pentesting puts autonomous AI systems in the driver’s seat. These agents don’t just scan for known signatures. They plan, adapt, chain exploits, and confirm impact much as an experienced tester would. But they do it at machine speed and scale.
The concept gained fresh attention this week. The Hacker News examined how such systems prove real attack paths while highlighting gaps in speed and enterprise-wide coverage that still require human oversight. The piece arrived amid a flood of vendor announcements and research reports that show the technology moving from experiment to production tool.
Consider the numbers driving urgency. Check Point Research recorded organizations facing 1,968 attacks per week in 2025, an 18 percent jump from the prior year. CrowdStrike’s 2026 Threat Hunting Report found detection leads triggered by AI agents growing 2.5 times faster than human-triggered ones. Attackers weaponize new vulnerabilities in about five days on average. Defenders often take weeks longer to patch.
That mismatch explains why companies from startups to established security vendors now race to field their own agents. Depthfirst introduced its version on September 30. The platform’s autonomous agents attack running applications like “a team of elite security researchers,” according to the company’s announcement. They work with or without source code, confirm whether static scan findings or bug bounty reports can actually be exploited at runtime, and replay attacks after fixes to verify remediation.
“Agentic Pentesting is surprisingly good: it even found a vulnerability in a feature that had been part of Supabase for two years and had already gone through multiple rounds of pentesting,” Etienne Stalmans, security engineer at Supabase, said in the depthfirst announcement.
The integration with code security sets depthfirst apart. Agents draw on continuous context from the codebase, prior findings, pull requests, and threat models. They authenticate to targets, explore interfaces and APIs, retain evidence across steps, and form new hypotheses. Security teams can trigger tests on demand or embed them in CI/CD pipelines.
French startup Fleuret AI offers another flavor. Founded in January 2026 by Yanis Grigy and Augustin Ponsin, the company launched its product in June. It bets that the best defense against AI-driven attacks is another set of AI agents running continuous tests. “What happens when the hacker is an AI agent? Fleuret AI is betting the answer is another AI agent,” wrote the French Tech Journal on October 6.
Larger players have joined the fray. Invicti launched Agentic Pentest in July, blending autonomous AI reasoning with its proof-based dynamic application security testing. “The future of application security isn’t about using more AI. It’s about using AI more intelligently,” Neil Roseman, CEO of Invicti Security, said in the company’s release.
Pentera earned recognition in Wavestone’s September landscape review of 69 agentic solutions. The consulting firm placed the company among the most promising for red team capabilities that demonstrate real control, stealth, and adaptability. Only 13 percent of reviewed tools met that bar. Wavestone expects these systems to augment auditors with faster reconnaissance, broader coverage, and automated re-testing. Yet it warned that technical controls, not just prompts, must keep agents inside approved boundaries.
Research backs the potential. Stanford’s RegLab pitted six AI agents against 10 professional penetration testers on a live university network of roughly 8,000 hosts. The ARTEMIS framework finished second overall, beating nine of the human participants. Such results fuel optimism. But they also raise hard questions about reliability, safety, and what happens when things go wrong.
The Institute for AI Policy and Strategy sounded a cautionary note in its September 29 report. Offensive cyber agents capable of autonomous attacks represent an emerging national security threat. In August 2026, OpenAI disclosed that an unreleased model had reached the company’s internal “critical cyber capability threshold.” That benchmark covers models able to devise and execute novel end-to-end strategies against hardened targets from high-level goals alone.
“Offensive cyber agents that can autonomously execute cyberattacks are an emerging national security threat,” the report’s authors, Christopher Covino, Jam Kraprayoon, and Matthew Mittelsteadt, wrote. They argue AI-enabled pentesting and red teaming could serve as a counterweight, acting as a force multiplier that identifies vulnerabilities faster, lowers assessment costs, and matches the speed of offensive agents.
Yet deployment carries risks. High-autonomy testing against production systems could disrupt operations. A development-to-deployment gap still limits how quickly defenders can adopt the latest models. Governance matters. Vendors now emphasize scope enforcement, blast-radius controls, independent validation of findings, and audit trails.
Kroll ran an agentic AI-assisted assessment for a banking client worried about frontier models such as Mythos. The consultant-led engagement used a model-agnostic harness that gave agents access to the same tools human testers employ. It uncovered high-impact risks while keeping professional judgment in the loop. The exercise showed the value of hybrid approaches that combine AI scale with human precision.
Debate continues over exactly how autonomous these systems should be. Barrion’s analysis draws a clear line. Traditional scanners follow fixed rules. AI-assisted testing offers suggestions to humans. True agentic pentesting lets the agent decide the next step inside approved boundaries, plan attacks, adapt, and chain exploits. Even then, proof of exploit remains essential because agents “can be confidently wrong.”
OWASP’s draft Autonomous Penetration Testing Standard attempts to bring order. Its four autonomy levels range from assisted, where humans command every action, to fully autonomous systems that manage multi-target campaigns with adaptive strategy and only periodic review. The standard also spells out requirements for scope enforcement, safety controls, human oversight, and reporting.
Enterprise buyers now ask tougher questions. Can the tool handle business logic flaws that depend on context no model fully grasps? Does it produce evidence strong enough for compliance teams? What happens when an agent follows a redirect outside the target scope or consumes excessive resources? Cost falls as models grow cheaper, yet quality and governance still separate credible offerings from marketing hype.
Early adopters report meaningful gains. Some organizations run tests after every major code change rather than once or twice a year. They revalidate fixes automatically. Human experts focus on novel attack paths, complex applications, and decisions that require deep business understanding. The result looks less like replacement and more like a new division of labor.
But the technology remains immature. No single vendor dominates. Architectures evolve quickly. Open-source projects multiply alongside well-funded startups such as XBOW, which raised $120 million in March after topping a HackerOne leaderboard. Incumbents like Pentera and Invicti push hybrid models that blend deterministic engines with reasoning agents.
Policy makers watch closely. If offensive agents can compromise hardened targets with minimal human direction, defenders must match that capability without introducing new risks to their own environments. The Institute for AI Policy and Strategy calls for piloting AI-enabled red teaming in controlled settings, developing safety standards, and closing the gap between research breakthroughs and safe operational use.
One thing looks clear. The era of the annual pentest as primary defense is ending. Continuous, trigger-driven validation aligned to actual risk exposure is taking its place. Agentic systems will form a core part of that shift. They won’t eliminate the need for skilled humans. Instead they amplify what those humans can achieve.
Security leaders evaluating these tools should demand more than slick demos. Look for independent validation of exploits, technical guardrails beyond prompt engineering, clear audit trails, and evidence the system handles real-world complexity rather than just synthetic benchmarks. The agents are here. The question is whether organizations will govern them wisely before adversaries do.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | AI Agents Slip the Leash: How Frontier Labs Lost Control of Their Own Creations | 0 | 8.59 | 02-10-2026 |
| 2 | When AI Agents Turn on Their Masters: Hackers Lose Email Harvest to Rogue Security Tools | 0 | 8.35 | 02-10-2026 |
| 3 | The Sandbox Illusion: Why AI Agents Escape and How to Spot Real Containment | 0 | 11.4 | 08-10-2026 |
| 4 | Webinar: How to Govern AI Agents, Reduce Excessive Access, and Control Shadow AI | 0 | 9.05 | 28-09-2026 |
| 5 | AI Agents Promise Help but Deliver Havoc: Inside the Push for Real Rules | 0 | 11.06 | 03-10-2026 |
| 6 | The SOC Doesn't Need to Start Over with Every Alert | 0 | 8.65 | 25-09-2026 |
| 7 | AI Cybersecurity Threats: Intelligence vs. Authority | 0 | 6.33 | 29-09-2026 |
| 8 | When AI agents swarm, can banks keep up? | 0 | 10.38 | 30-09-2026 |
| 9 | Persona Expands Identity Theft Protection Tools as New Data Reveals Identity Fraud Is Becoming More Adaptive and Sophisticated | 0 | 9.13 | 01-10-2026 |
| 10 | Cisco Integrates Agentic AI into Webex to Automate Enterprise Workflows | 0 | 8.31 | 07-10-2026 |