Skip to main content
Back to Blogs
Guide8 min read

What Is a Good Score on a Psychometric Test? Percentiles, Stens and Cut Scores

Raw scores mean almost nothing on a psychometric test. How percentiles, stens, stanines and norm groups actually work, and why the same performance can be a pass at one employer and a fail at another.

Almost every question candidates ask after an assessment is a version of the same one. I got 32 out of 40, is that good? I finished 28 of 50, is that a fail? The honest answer is that nobody, including the employer, can tell you from that number alone, because the raw score is the first thing the scoring process throws away.

Why your raw score is discarded

A raw score is meaningless on its own for a reason that becomes obvious once you see it: it depends entirely on how hard the test was. Twenty-eight correct out of fifty is an excellent result on a hard test and a poor one on an easy test, and the candidate has no way of knowing which they sat.

Test publishers solve this by comparing your performance against a group of people who have already sat the same instrument. That group is called the norm group, and the comparison is what gets reported. Your raw score exists for about one step of the calculation and then disappears.

Two consequences follow immediately, and they surprise people. First, there is usually no percentage anywhere in an assessment report. Second, the same raw score can produce different reported results for two candidates, if they were compared against different groups.

The norm group decides everything

This is the single most important idea on this page. A norm group is the reference sample your result is scored against, and publishers maintain several for the same test.

General population
Everyone. The most generous comparison available, because it includes people who would never apply for the role. A result at the 70th percentile against the general population can be an unremarkable result against graduates.
Graduate or professional
People with degrees, or people in professional roles. The standard comparison for graduate scheme testing, and a much tougher one. This is why a candidate who did well at school can land mid-range here without having got worse at anything.
Role or sector specific
Applicants to a similar role, or people already doing it. The most informative comparison and the least common, because it needs enough data from that specific population.
Country or language
Separate norms per market. Relevant if you are sitting an English-language test outside an English-speaking country, and one reason results are not portable between regions.
The question worth asking
If you are ever given feedback on an assessment, the useful question is not “what did I score?” but “which norm group was I compared against?”. A 55th percentile against a professional norm group and a 55th percentile against the general population describe genuinely different performances.

Percentiles, stens, stanines and T scores

Four reporting scales dominate, and they all express the same underlying comparison with different granularity. Learning to read them takes about two minutes and saves a lot of misinterpretation.

Most common in the UK
Percentile
The percentage of the norm group you scored above. The 70th percentile means you outperformed 70 per cent of that comparison group. It is not a percentage of questions correct, and confusing the two is the most frequent scoring error candidates make.
Common in the Netherlands
Sten score
A one to ten scale, where 5 and 6 are the middle band containing most people, 1 to 3 is below average and 8 to 10 is well above. A sten of 7 is a good result, not a mediocre one, because the scale is not out of ten in the school sense.
Common in older reports
Stanine
A one to nine scale with 5 at the centre. Same idea as a sten with a different number of bands. Occasionally used in education and public sector testing.
Common in research reports
T score
Centred on 50 with a standard deviation of 10, so 60 is one standard deviation above the mean, roughly the 84th percentile. Turns up in personality reporting more often than in ability reporting.

The rough conversions are worth memorising, because reports frequently mix scales. A sten of 7 is around the 77th percentile. A stanine of 7 is around the 89th. A T score of 60 is around the 84th. If a report gives you a band rather than a number, that is deliberate: publishers band results to stop people over-reading small differences that are within measurement error.

Cut scores, and why nobody will tell you the pass mark

A cut score is the threshold an employer applies. Three things about it explain most of the confusion around assessment results.

The employer sets it, not the publisher

Test publishers supply the scale. Employers decide where to draw the line, and they do it against their own hiring needs. This is why a page claiming “you need the 70th percentile to pass at company X” should be treated as unsourced folklore.

It moves between intakes

A cut score is often set to produce a manageable number of candidates for the next stage. In a year with twice the applications, the same performance can fall below a line it cleared last year. Nothing about you changed.

Sometimes there is no cut score at all

Plenty of processes use assessment results as one input among several, or as a ranking rather than a gate. In those processes the question “did I pass?” has no answer, which is why recruiters sometimes seem to be dodging it.

Why an adaptive test feels hard

An adaptive test selects your next item based on how you answered the last one. Get one right, get a harder one. Get one wrong, get an easier one. Over a sitting this converges on the difficulty level where you are answering roughly half the items correctly.

Which means: on a well-functioning adaptive test, everybody finds it hard, and finding it hard tells you nothing about your result. A candidate who breezed through an adaptive test has usually been converged downwards. A candidate who left feeling battered by the last five items has usually been converged upwards.

Adaptive designs also break the arithmetic candidates try to do afterwards. Your item count is not comparable with anybody else's, because you did not sit the same items. Aon, Korn Ferry Talent Q, Assessio's Adaptive Matrigma and Sova's reasoning tests all use adaptive or semi-adaptive designs, and practising them is largely about getting used to that feeling.

What you can actually find out

  • Ask for your report. In the UK and EU you generally have a right to personal data held about you, and in several markets that includes assessment reports. Publishers write feedback reports precisely so they can be shared.
  • Ask which norm group was used. A recruiter often will not know, but the question sometimes reaches someone who does, and it is the only number that makes a percentile interpretable.
  • Do not ask for the pass mark. Employers treat cut scores as commercially and legally sensitive, and asking rarely gets an answer.
  • Do not compare with friends. Different norm groups, often different item sets, sometimes different tests with the same name. The comparison is not measuring anything.

The practical upshot

Stop trying to reconstruct your score. It is not recoverable from what you remember, and even if it were, the raw number is not what anyone is looking at. The controllable things are the ones that come before the sitting: knowing the format, working at a steady rate, and knowing whether the test you are on penalises guessing. Those change your percentile. Post-hoc arithmetic does not.

If you want to know how much of the outcome preparation actually moves, the companion piece on whether practising psychometric tests works gives the honest version.

Ready to Start Practicing?

Apply these strategies with our comprehensive practice platform

Start Practising Free