The methodology
How we test
“Best AI pentester” isn't one skill. A good tester finds a lot of the surface and can chain a foothold into an objective. Those are different abilities — a model can be strong at one and weak at the other — so we measure them separately, on separate labs, and never blur them into a single blended score.
What one run is: one autonomous session
A run is one autonomous engagement — the model is dropped in a locked, disposable sandbox with a target and a scope (no hints, no answer key) and does its own recon, then hunts and exploits with its own tools, all in one pass. Because results move run to run, we run each model three times per lab (R1 · R2 · R3) and report the mean (with a floor), not a lucky best. A run that ends abnormally is thrown out and re-run. That honesty machinery is shared by both axes.
Coverage — did it find them?
The coverage lab is a multi-service enterprise app — real endpoints, real misconfigurations — seeded with the kind of vulnerabilities we've hit on engagements. We score two things: recall (share of planted vulns found, averaged over the three runs) and precision (signal vs noise; a planted decoy reported as real costs here), multiplied together. Every match is mechanical — against a ground truth we verified at the source-codelevel, never an opinion. We report the mean across runs and a floor alongside it, so one lucky pass can't carry the score.
Coverage has a ceiling, though: counting one-shot vulnerabilities rewards being systematic, not being clever. That's why it's only half the story.
Exploitation — could it actually get in?
The exploitation lab is a realistic app whose weaknesses chain. The objective — a shell, admin, the data — is reachable only by stringing several steps together: an SSRF that pivots to an internal service, credentials reused from one hop to unlock the next, a business-logic flaw that a scanner will never flag. We don't score how many bugs you named; we score how far along the chain you got.
Each step of each chain hides a unique marker— a secret string an agent can only obtain by actually completing that step (the random suffix means it can't be guessed, only recovered). A model “reached milestone 3” if, and only if, milestone 3's marker appears in its output. That makes depth as objective as a code-verified vuln match — no self-grading, no partial credit for a good story.
Why depth is the harder, more honest signal
Difficulty is baked into the weights: shelling the easy chain is worth little; reaching the deep end of the hard one is worth a lot. The final objective on the hardest chain requires synthesis— connecting a secret leaked at the very first step to a lock at the very last — which even strong models miss. So the score has headroom: as models improve they climb the milestones, and the top of the chain stays aspirational. The failure mode becomes “couldn't reason through it,” not “didn't run enough.”
Two labs aren't a population
Two labs is a start, not an end. A model strong on these stacks may be weak on others, and small samples carry real variance — we show it. More labs are how this grows: each new lab is another target, another ground truth, and one of the two scorers. The leaderboard reports each lab and a simple average, the way a transcript reports subjects and a GPA — never one number that buries where a model is weak.
Related: how we score models, out of 1000 · keeping the benchmark honest · the target & ground truth.
← Back to the blog