What Score Estimates Can and Cannot Tell You

Every prep tool shows you a projected score. Most of them are more optimistic than they should be, and understanding why will change how you read your own.

The MCAT is scored 472 to 528, with 500 near the middle of the distribution. The temptation when building an estimator is to map percent-correct straight onto that range: 50% correct becomes 500, 80% correct becomes roughly 517. It is simple, it is intuitive, and it is wrong in a way that consistently flatters the student.

Percent correct is not linear

On real material, the relationship between raw accuracy and scaled score is a curve, not a line. Around half correct sits near the middle of the distribution. The top of the scale requires something close to perfection, and the gap between 85% and 95% is worth far more scaled points than the gap between 45% and 55%. An estimator that treats those as equal will overstate strong students slightly and weak students substantially.

Not all practice questions predict equally

The larger problem is what the questions are measuring. A definitional flashcard — "the ___ states that no two electrons share the same four quantum numbers" — tests recall. The MCAT rarely tests recall directly. It tests whether you can apply a principle to a situation you have not seen, read an experiment critically, or interpret data.

Those are different skills, and they come apart. It is entirely possible to score 85% on recall and 55% on applied reasoning. A student in that position who sees a single blended number is being told something misleading, and the direction of the error is the dangerous one: they will believe they are ready.

A defensible estimate therefore weights applied questions far more heavily than recall, and says so. If your estimate is built mostly on flashcards, it should be labelled as provisional rather than presented as a prediction.

Sample size, and the case for a range

Twenty questions tell you very little. Accuracy on small samples swings widely for reasons that have nothing to do with knowledge — which topics happened to come up, whether you were tired, whether two questions happened to be ones you had seen.

Any estimate should widen its uncertainty when the sample is small, and should show a range rather than a single number. A tool that displays "512" after thirty questions is expressing a confidence the data cannot support. "505 ± 9" is less satisfying and considerably more honest.

Why the number matters less than the breakdown

The projected score is the least actionable thing a practice tool can show you. It tells you where you stand and nothing about what to do. The useful information sits underneath it:

Reading your own numbers

A few rules that hold regardless of which tool you use:

None of this makes practice estimates useless. Tracked over months, the trend is informative even when the absolute number is not. The mistake is treating a proxy as a prediction, and then planning a test date around it.