A single number with no interval around it is a guess wearing a lab coat.
The question everyone actually wants answered is not "how am I doing." It is "is one more practice test worth it, or should I book the exam." That is a question about confidence, and almost every practice tool answers it with a single number carrying no confidence at all.
Suppose you sit twenty questions and get fourteen right. What does that tell you about a 53-question exam spanning twenty-four objectives?
Less than it appears. Twenty questions cannot cover twenty-four objectives, so several were not sampled at all and your score says nothing whatsoever about them. Of the ones that were, most got one or two questions. Getting one question right is not 100% mastery of an objective; it is one question. The headline figure is real, but the thing it is a measurement of is mostly still unknown.
This is why an estimate needs an interval, and why the interval has to be driven by how much evidence exists rather than by how much your scores have bounced around. Two sittings that happen to agree are not more informative than one; they are just two thin samples pointing the same way.
A readiness estimate should start wide and narrow as it earns the right to. Ours opens at roughly ±60 scaled points when it knows nothing about you, and closes toward ±12 once every objective has been sampled properly — treating about ten items per objective as the point where an objective is pinned down rather than guessed at.
The interval, not the midpoint, is what answers the real question. Three verdicts fall out of it:
That third verdict is the one most tools will not give you, because it looks like the product failing to answer. It is the opposite. "Too close to call, and here is what would settle it" is the answer that changes what you do today.
There is a fourth state, and it matters more than it sounds: not enough data yet.
Answer four questions and any model can still produce a number. It will be mostly prior — a statement about the model's assumptions rather than about you. Reporting "below the cut" to someone who has answered almost nothing is a confident claim drawn from no evidence. It is wrong on the statistics, and it is the worst possible first thing to show a person who has just started.
So below a threshold of evidence the number should be withheld and the tool should say what it needs instead. A readiness estimate that refuses to guess is worth more than one that always answers, for the same reason a lifetime accuracy percentage is worth less than it looks: the confident number is the easy one to produce and the hard one to trust.
If you want to know whether you are ready, the sequence is short. Get every objective sampled — not just the ones you enjoy, and not just the heavy ones. Recombine the results by blueprint weight rather than averaging them flat. Then read the interval rather than the midpoint, and keep practising only while 720 is still inside it.
Stop when the bottom of the interval clears the cut. That is the whole answer, and it is knowable well before you have exhausted a question bank.
ExamGauge is built on the argument above: 1523 original practice items across four Claude certification exams, scored on the real 100–1000 scale against the 720 cut. The diagnostic is free and needs no card.