ExamGauge

Your practice percentage is measuring the wrong thing

Any tool that re-asks the questions you got wrong poisons its own accuracy statistic. Here is how far off it gets, measured.

Every adaptive practice tool does the same useful thing: it brings back the questions you got wrong and quietly retires the ones you have nailed. That is the entire point. It is also what makes the accuracy percentage it shows you close to meaningless.

The record such a tool leaves behind is not a fair sample of your ability. It is deliberately enriched with the questions you were worst at. Reading a lifetime accuracy figure out of it tells you about the selection algorithm, not about you — and the distortion gets worse the better the algorithm works, because a better algorithm re-shows your weak items more aggressively.

How far off it gets

This is measurable rather than arguable. On the eighteen-month record this engine was built from, across the Developer bank:

MeasureReads
Impressions recorded1,493
Distinct items behind them299
Average times each item was shown4.99
Lifetime accuracy (every answer ever given)34.8%
Current state (each item once, as it stands now)62.5%

Those two numbers describe the same person on the same day. The first counts a hard question you eventually mastered as five failures and one success. The second counts it as one item, currently correct.

Converted onto the scale the real exam reports, the gap is worth roughly 300 scaled points. Lifetime accuracy would have told a candidate who was ready to sit that they were nowhere near the cut.

The distortion scales with re-showing

The same measurement on the Architect bank, where items had been shown 1.65 times each rather than 4.99, reads 64.8% lifetime against 70.9% current state — a six-point gap instead of a twenty-eight-point one.

That relationship is the tell. If your practice tool never repeats anything, its percentage is roughly honest and roughly useless, because it never revisits what you got wrong. If it does repeat, the percentage decays as a measure of you and improves as a measure of how hard it is working. There is no setting at which one number does both jobs.

What to use instead

Let each distinct item vote once, at its current state. That is the whole correction: not how many times you have answered something, but whether you had it right the last time you saw it. Re-showing a hard item ten times then stops registering as ten failures.

Two smaller things matter alongside it. Thin evidence should read as uncertainty rather than as a score — one right answer out of one is not 100% mastery, and an objective you have never sat is unknown, not zero. And the per-objective results have to be recombined by blueprint weight before they mean anything, because a drill that deliberately over-samples your weak areas produces a raw score that is pessimistic by construction.

None of that is exotic. It is just what you have to do once you accept that the practice record was never a random sample — and that the tool spoiled it on purpose, for good reasons.

ExamGauge is built on the argument above: 1523 original practice items across four Claude certification exams, scored on the real 100–1000 scale against the 720 cut. The diagnostic is free and needs no card.

More on this