Technical blog
Everything behind the leaderboard, in detail: how a test runs, the targets and how their answer keys are built, the engine, and how a model is scored — with the diagrams, the real numbers, and the limits. Enough to decide whether to trust the board.
Two skills, measured separately: coverage (how much a model finds) and exploitation (how far it chains a foothold into an objective). How a test runs, and how depth is verified with planted markers.
Why we score against a private, code-verified target instead of public CTFs — and why building the answer key from source code, not the README, is what makes the score mean anything.
One autonomous agent in a locked sandbox, and a hard rule: nothing is scored unless it can be proven — and matched against a source-verified answer key.
Recall, consistency and precision — what each measures, why cost is excluded, how a private answer key defeats contamination, and where this benchmark's limits are.
The scoring bugs we found in our own harness and fixed: excluding truncated scans, refusing to score consistency from one run, measuring cost honestly, and never penalising a model for how it formats its findings.