We Scored 99% on AI Pentesting Benchmarks by Turning Our Agents Into Something They’re Not

When we built Cobalt Autonomous Pentest, we knew we had to provide an independent assessment of our quality so you didn't just take our word for it. There are many benchmarks out there, but few that are really comprehensive. In fact, as soon as one gets decent, it's used to train new models, so it can’t be fairly evaluated. We played with several benchmarks and ended up zeroing in on the XBEN Validation Benchmarks. We hit 103 out of 104 on XBEN, but this score doesn’t actually matter.

What do you mean the benchmarks don’t matter?

Benchmarks reward a very focused, vertical slice of what a real pentester looks for in a web application pentest. That simply isn’t what makes a real penetration test valuable. Our agents are tuned to use our tried and proven methodologies based on our 13 years of experience conducting pentests. They aren’t built to compete in a capture-the-flag exercise. In order to hit that 99% coverage, we had to modify the agents’ win conditions. We had to make them something they aren’t.

These benchmarks are dockerized, Jeopardy-style CTF challenges with one synthetic lab-grown vulnerability per challenge, each mapped to an OWASP Top 10 class. Finding planted static text flags inside unhardened micro-containers is the minimum bar for a robust offensive security tool, not the finish line. Real, enterprise-grade applications operate on multi-tenant RBAC, with federated identity providers, and stateful business logic workflows behind WAFs. Manufactured benchmarks sidestep all of that which means role boundaries and access control bypass capabilities are never assessed. Access control is the second most commonly found vulnerability in our web applications pentests (2026 State of Pentesting Report), so the gap here is significant. While more advanced than scanner-style benchmarks, these CTF benchmarks are aging fast.

Anytime new data is released and new challenges are built, frontier models consume them as training data. Likewise, intentionally vulnerable applications with answer keys are susceptible to this same issue. Running evals against your agent’s harness gets less and less valuable when the answers are trained into the models themselves.

Yes, we gamed the system. Then we shelved that version

We built a custom configuration for our agents to hit 99% on these benchmarks, one that will never see the light of day in our product because it isn’t what security teams need. We tuned an instance of our autonomous pentest to XBEN and several known vulnerable applications like OWASP Juice Shop and Broken Crystals. While we got increasingly higher scores on these benchmarks by tweaking our prompts, and modifying the agent’s specializations, we saw regressions in how the agents performed on real applications. Tuning to these benchmarks in reality made our agents less effective penetration testers.

willa image

So how do we measure autonomous pentesting solutions?

Where do we go from here, knowing that the benchmarks are cooked, the results don’t translate to real business outcomes for security teams, and the models are cheating by ingesting the answer keys? We have to find a better way to measure outcomes for these agents.

What makes a pentest effective is not expecting that there is only one answer to every challenge. It's understanding the risk to your business. We test complex vulnerability classes that CTF benchmarks ignore completely, yet make up key steps in common attack chains. This includes vulnerabilities like cross-user insecure direct object reference (IDOR), administrative middleware bypasses, Next.js build topology disclosures, and federated identity checks.

We have to measure the agents’ performance against human creativity, ingenuity, and judgment. Cobalt is uniquely positioned to meet this need.

Where Cobalt Autonomous Pentest shines

When you build an autonomous pentesting product after years of doing human-led pentests, you measure differently. You stop asking 'did it find the flag' and start asking 'would I trust this report if I were the one receiving it.' We run more than 5,000 pentests a year at Cobalt, and over a decade of engagements showed us exactly what the agents needed to do to earn that trust.

In a standard run, Cobalt Autonomous Pentest enumerates over 100 unique endpoints, maps 400-plus parameterized routes across multiple user role contexts, and fires more than 10,000 targeted HTTP probes without knocking assets offline. Those numbers matter because they reflect what a thorough pentest actually covers, not what a benchmark rewards.

When running against business applications and comparing the results with human-led engagements, the outcomes were fascinating. The findings did not overlap the way we expected. Each approach surfaced issues the other missed. We will walk through what we found, including specific finding types and where each approach had the edge, in our live virtual demo on August 27th. But our takeaway was that it's the combination of the two that is the most powerful. Humans and AI. Working together.

willa image 2

This comparison between autonomous and human-led engagements shaped a core product decision. As someone who has spent years on the pentesting side, I know what happens when automation runs without direction: it covers ground fast but misses context. So we built Cobalt Autonomous Pentest with human oversight. Elite Cobalt pentesters review the execution plan, enforce scope and methodology, and ensure the AI operates within the boundaries your program requires. Not because automation needs a babysitter, but because an AI pentest without expert judgment is just a scan with better marketing.

What this actually means for our customers

Autonomous pentesting is best for speed, breadth, cost, and frequency. Human-led pentesting is best for depth, creativity, and judgment. Companies that ask you to pick one or the other are only selling you the one they have. The good news is that you don’t have to choose.

A 99% benchmark score tells you how agents perform in a vacuum. It does not tell you whether your business is protected. Let us show you how Cobalt Autonomous Pentest can do that for you. Join us for a live virtual demo on August 27th with the Cobalt product team.

 

Back to Blog
About Willa Riggins
With over 20 years of hands-on experience in application development, information security, and communications, Willa Riggins, Sr. Staff Product Manager - Tester Team, brings a unique, holistic perspective to the entire software development lifecycle. She's a recognized speaker at industry events like DEF CON and BSides, and a community leader who has shaped various local hacker and security groups. More By Willa Riggins