Today, we launched Cobalt Autonomous Pentest, a new way to run offensive security testing across your entire application portfolio. It pairs AI testing with elite Cobalt pentesters who direct every engagement. Launch a pentest in minutes and get exploitable findings in 24 hours.
I have spent a quarter century running offensive security operations, and one thing has not changed in all that time. Some applications get pentested. Most do not.
It's a simple math problem. A security team and an annual budget that are finite, and a portfolio that grows every sprint. The critical applications get proper, thorough testing. Everything else gets a scanner, or gets nothing. That was a tolerable problem when everyone moved at the same speed, but it’s become untenable.
AI changed both sides of the equation at the same time. Engineering adopted AI and started shipping faster than any security team can test, even with the best DevSecOps practices in place. Attackers adopted AI too, and now run reconnaissance and exploitation at machine speed against everything you have exposed. The portfolio you were not testing is now where attackers are acting.
The trap security teams are in
Security teams need AI in the testing loop. However, they don’t fully trust AI yet as the technology is still evolving. In fact, according to Omdia’s Research Survey, 94% of security teams keep humans in the loop on AI agents. The instinct behind that number is right; security needs oversight to balance out the immaturity. The answer is not less AI. It is AI that a human directs.
Introducing Cobalt Autonomous Pentest
Cobalt Autonomous Pentest runs AI-driven offensive security across your applications, with a Cobalt pentester overseeing every engagement. The AI does the testing. The pentester inspects the scope, leveraging our testing methodologies, and stands behind how the engagement runs.
You launch in minutes, and get findings in 24 hours. And because a human pentester manages the work, you get results you can hand to a developer without the caveat that an unvetted AI scan requires.
How it works
An autonomous pentest is not a fancy scanner—it reasons the way an attacker does. It maps the application, selects tools, chains steps together, and attempts exploitation. When a path works, you get proof: the details, the proof of concept, the steps to reproduce it. Not a theoretical finding a developer will dispute.
The engine is model-agnostic. It orchestrates the best models and tools for each phase of the work and updates as the frontier moves, so your testing keeps pace with the state of art in pentesting methodologies and adversarial TTPs. Lock a security tool to a single model and it falls behind the moment the next one ships.
Cobalt Autonomous Pentest is built on 13 years of real pentest data, the largest dataset of its kind, crafted with the same methodologies, tactics, and techniques leveraged in engagements today and known to deliver results. It knows what attackers actually do, not known issues scraped from the internet or how a CTF was designed.

A Cobalt pentester reviews the plan the AI creates, checks that the scope and boundaries are enforced, ensures full methodology coverage, and authorizes the testing that follows. The AI executes inside the boundaries the pentester sets.
So, why do we feel this strongly about leaving a human in the loop? Autonomous action against a production system with no one directing it, and no one accountable for where it goes, is an unbounded risk. Someone has to own what the test does and be able to answer for it.
What benchmarks fail to tell you about quality
Quality is the first question anyone asks, and it is a hard one to answer from the outside. You can’t really know how good an autonomous pentest is until you run it against your own applications. We understand the need for a proxy, so we did what everyone else does. We ran a benchmark. Early iterations of our testing passed 100% of the XBOW validation benchmarks.
I am glad we scored so well, but we found that doing well at contrived challenges did not reflect the reality of customer web applications. These testing suites are built as capture-the-flag puzzles, and if you optimize your agents to win them, that is what you get: an agent good at CTFs. These are not Ender’s Game war simulations. These are sudoku. A pentest is not a CTF. A real application does not have a known answer waiting at the end.
Our quality comes from somewhere else, built on our depth of experience. We have data and accumulated experiences from 13 years of pentesting real applications, with an understanding of what makes a good test, and hard-won knowledge of which tools and techniques actually find vulnerabilities across 5,000-plus pentests a year. Behind that sits the Cobalt Core, our elite global community of 500+ rigorously vetted pentesting experts. These professionals average 11 years of experience and hold top certifications like OSWE, OSCP, and CREST.
That is why humans are at the center of everything we do. The attacker is still a person, just armed with capable AI now. So we mirrored that. A white hat with AI in hand, finding the risks in your applications before someone less friendly does.
Where it fits in your strategy
I am excited about our new Autonomous Pentest offering, but I want to be clear about what it is not. It does not replace human-led, AI-powered pentesting. It complements it. Cobalt Autonomous Pentest brings offensive testing to the applications that pose real risk but are not getting tested today. Human-led pentesting still does the deep work, with manual expertise, compliance-grade engagements, and asset types that demand the highest precision. This includes the AI and LLM applications where Cobalt has led the industry and where the risk to organizations keeps climbing.
Cobalt Autonomous Pentest is built to strengthen an offensive security program, not to stand alone. Combine Autonomous Pentest with human-led pentesting, red teaming, and DAST, all delivered through the Cobalt Offensive Security Platform, which provides real-time collaboration with testers, calendar planning, analytics and benchmarking, as well as seamless integration with the tools your teams already use. Build a program that validates real risk continuously across the full attack surface, and adjust the mix as your priorities change.
I have watched offensive security go from a specialist craft to an established discipline. And the industry is at another turning point. Continuous testing across the whole portfolio is becoming the baseline, the way continuous integration did for development. The programs that make that shift will run on automation, AI, and expert human direction together. That is the program we are building toward, and Cobalt Autonomous Pentest is a piece that makes it possible.
Come see Cobalt Autonomous Pentest in action at Black Hat USA, booth 4903.