ExamGauge

A 720 is not 72%

Scaled scores and weighted blueprints both break the arithmetic people do in their heads.

The Claude certification exams report a scaled score between 100 and 1000, and the cut is 720. It is tempting to read that as 72% and work backwards: fifty-three questions, so thirty-eight right and I am through. Both halves of that arithmetic are wrong, and they are wrong in opposite directions, which is why the mistake survives contact with practice results.

A scaled score is not a percentage

Scaling exists so that scores mean the same thing across different forms of an exam. Sit a slightly harder set of questions and the raw mark needed for a given scaled score comes down; sit an easier one and it goes up. The scaled number is the constant, and the raw percentage behind it moves.

So there is no fixed number of questions that equals 720. Anyone who tells you exactly how many you need to get right is describing a form they cannot have seen. The honest version is directional: a scaled cut of 720 on a 100–1000 range sits somewhere near, but not at, 70% raw — and the exact figure is not yours to know or to plan around.

What follows practically is smaller than it sounds. Stop counting questions. The quantity that survives across forms is the proportion of the blueprint you have covered, not the proportion of one practice set you happened to answer.

The weighting breaks it a second time

The larger error is treating every question as worth the same. On a blueprint-weighted exam they are not, and the spread is not subtle.

On CCDV-F, the published domain weights run from 33.1% down to 2.6% across 8 domains. Applications and Integration is 33% of the paper on its own — about 18 of its 53 questions. Eval, Testing, and Debugging is 2.6%, which is 1.

Now take two candidates who both score 70% on a flat practice test. The first dropped their marks in Eval, Testing, and Debugging. The second dropped the same number in Applications and Integration. They have identical practice percentages and materially different chances of passing, and no flat score will ever tell them apart.

This is what a percentage hides and what a weighted estimate surfaces. A practice result is only informative once each objective's result is multiplied by that objective's share of the real exam and the pieces are added back up. Which also means a drill that over-samples your weak objectives — as any adaptive drill should — produces a raw score that is pessimistic by construction and has to be corrected before it is shown to anyone.

What the number should look like

A useful readiness figure is on the same 100–1000 scale as the exam, so it is directly comparable to a real score report rather than to a percentage that means nothing outside the tool that produced it. It is weighted by the published blueprint. And it comes with an interval, because an estimate from four questions and an estimate from four hundred should not look alike.

You can see the weights we score against for each exam on the blueprint pages: CCAR-F, CCDV-F, CCAO-F and CCAR-P. They are Anthropic's published numbers, not ours — which is exactly why they are worth planning against.

ExamGauge is built on the argument above: 1523 original practice items across four Claude certification exams, scored on the real 100–1000 scale against the 720 cut. The diagnostic is free and needs no card.

More on this