Automated, Agentic, Autonomous: The AI Pentesting Vocabulary Problem

As global penetration testing undergoes its biggest transformation since the invention of PTaaS, offensive security practitioners and buyers are being flooded with marketing jargon and technical comparisons that blur the differences. Prefixes like "Automated," "Agentic," and "Autonomous" are increasingly used interchangeably - and often randomly - to describe features, capabilities, and aspirations of newly AI-powered pentesting products and services. The problem isn't just aesthetic. When vocabulary loses meaning, buyers can't make informed decisions, and vendors have no incentive to be honest about what their products actually do.

Before evaluating any AI pentesting claim, it helps to establish what pentesting actually is and how much automation and intelligence it already contained before AI arrived. The 2026 State of Pentesting Report is worth reading for data on the industry's current state, but first, we need to solve the vocabulary problem.

Modern Pentesting Was Already Heavily Automated

Much of the marketing hype around AI-powered pentesting anchors itself around comparisons with "traditional" consultant-delivered services as if the baseline is a lone analyst with a notepad and a Kali Linux VM. That picture is a couple of decades out of date.

In the 35-year history of commercial pentesting, there has been significant evolution. While modern pentesting practices are tied to rigorous and detailed testing methodologies (which is a very good thing), they're also extensively tool-based and automated. Across the cybersecurity industry, standard web application pentesting is over 80% automated through a mix of off-the-shelf and self-assembled tools that the pentester develops based on their needs throughout their career. Even IoT device reverse engineering makes extensive use of automation, often accounting for 50%+ of testing activities and frequently uncovering 90%+ of the actual security findings.

Over the last two decades, those tools have transformed pentesting from an art into a science. Year on year, they have measurably evolved. Today, pentesters employ a core set of methodology-following, target-specific tools packed with their own intelligence. Many of these tools have had AI-powered reasoning and fuzzing capabilities for five or more years. The most popular tools now ship with embedded LLM and vendor-trained AI models for advanced vulnerability discovery and automated exploitation. As I wrote back in 2017, the gap between AI hype and AI capability in security was already a significant problem, and the intervening decade has closed some of that gap while opening new ones. All of this capability sits in the hands of skilled cybersecurity professionals, deployed daily across thousands of pentests.

The honest baseline, then, is not "human versus AI." It's "AI-augmented human versus AI operating without one." That distinction matters enormously.

The Human Is Still the Secret

AI, automation, and tooling, while critical components of modern pentesting, achieve little if not properly employed by trained professionals. Pentest quality and customer satisfaction are primarily in the hands of the consultants guiding the tools and conducting the work within an authorized scope. The human is the secret behind every good - and every ugly - pentest.

What separates a poor pentest from a great one isn't the toolset. It's the judgment applied to it.

Pentest quality spectrum
What separates a bad pentest from a great one?
Level Quality expectations
Bad An inappropriate selection of tools and configurations means scoped parameters are missed. Reporting is copy-paste tool findings littered with false positives and inactionable recommendations.
Poor Pentesting scope was met via tooling. Obvious false positives and duplicate findings are removed. Reported findings are limited to what the tools found, and everything looks and feels like copy-paste recommendations without proper target context.
Average Pentesting scope was met. Findings are almost exclusively limited to what the tools were capable of discovering and enumerating. False positives are removed, remaining findings are independently researched and enriched by the consultant, and manually validated where possible. Reports describe findings in context with specific fix recommendations.
Good Pentesting scope was met and refinements were made as the pentest progressed, in collaboration with the customer. Tool-based findings are supplemented with manually discovered vulnerabilities, independently researched, verified, and enriched. Reports describe all findings in context, carry specific fix recommendations, and incorporate risk calculations that include the threat landscape. Findings may be updated after customer fixes have been applied.
Great Everything that makes a Good pentest, and then some. New vulnerabilities and attack vectors are discovered that are clearly beyond what current tooling can uncover. Vulnerabilities are evaluated for exploitation and, if safe and with customer consent, are exploited — with layered defenses evaluated for pivoting opportunities. The customer receives options for remediation and disclosure of novel findings, and the consultant works closely with them to validate all fixes, including those in third-party affected products.

 

What Automated, Agentic, and Autonomous Actually Mean

Today, these three terms describe a spectrum of human involvement—or its absence—in the pentesting process across the rapidly evolving offensive security ecosystem. They are not synonyms.

Automated pentesting is what the industry has been doing for twenty years: tools executing predefined workflows against a target, with a human interpreting and acting on the results. The human remains central to planning, judgment, and reporting.

Agentic pentesting introduces AI systems capable of making decisions, selecting and chaining tools, and adjusting their approach based on what they find, but still operating within guardrails and, typically, with human review at key decision points. The human shifts from executor to overseer. The launch of Cobalt Autonomous Pentest earlier this year illustrates what this looks like in practice: continuous offensive security across an entire portfolio, with humans steering the most consequential decisions.

Autonomous pentesting is the full expression of AI-led testing: a system that scopes, executes, finds, validates, chains, and reports without meaningful human involvement. As of today, this exists in limited and targeted forms, and its quality ceiling sits well below what a skilled human-led team can deliver for most real-world targets.

These distinctions matter because the depth and reliability of findings - and the price - differ substantially across the three categories. Conflating them serves vendor marketing, not buyer decision-making.

AI is indeed ushering in a new era of pentesting. Buyers should expect that it will reduce the cost of certain categories of testing and broaden the assets that can be assessed within a given budget. As I've argued in The Cobalt Vision for a Human-Led, AI-Powered Future in Security Testing, the right frame is not replacement but symbiosis — and that framing starts with being clear about what each system can and cannot do. But the starting point for any honest evaluation is understanding where on that Bad-to-Great spectrum a given AI pentesting product can realistically operate — and for which targets, at which depth, and under whose oversight. The prefix alone tells you nothing.

Back to Blog
About Gunter Ollmann
Gunter Ollmann serves as Cobalt's Chief Technology Officer (CTO). With rich and diverse experience in cybersecurity innovation, Ollmann leads Cobalt's technology and services strategy, delivering AI-enabled offensive security solutions coupled with unmatched human ingenuity. More By Gunter Ollmann
5 Key Takeaways from the 2026 State of Pentesting Report
Pentest data reveals a stark divide between organizations with an ad-hoc testing model and those with a programmatic approach.
Blog
Apr 21, 2026