Technical blog

How the benchmark actually works

Everything behind the leaderboard, in detail: how a test runs, the targets and how their answer keys are built, the engine, and how a model is scored — with the diagrams, the real numbers, and the limits. Enough to decide whether to trust the board.

The methodology
How we test

Two skills, measured separately: coverage (how much a model finds) and exploitation (how far it chains a foothold into an objective). How a test runs, and how depth is verified with planted markers.

9 min · methodology · two axes
The lab
The target and the ground truth

Why we score against a private, code-verified target instead of public CTFs — and why building the answer key from source code, not the README, is what makes the score mean anything.

8 min · methodology
The engine
How the engine works

One autonomous agent in a locked sandbox, and a hard rule: nothing is scored unless it can be proven — and matched against a source-verified answer key.

9 min · architecture
The score
How we score models, out of 1000

Recall, consistency and precision — what each measures, why cost is excluded, how a private answer key defeats contamination, and where this benchmark's limits are.

10 min · scoring · limitations
The harness
Keeping the benchmark honest

The scoring bugs we found in our own harness and fixed: excluding truncated scans, refusing to score consistency from one run, measuring cost honestly, and never penalising a model for how it formats its findings.

8 min · integrity