iTestHub
Sign In
< All articles
Assessment science 3 min read

What makes a psychometric test valid? Reliability, validity and fairness explained

Rob Dominic, Chartered Psychologist · Published 23 Aug 2026 · Updated 23 Aug 2026

Three words get used interchangeably in assessment marketing and mean quite different things. Understanding the difference is the difference between buying a test and buying a claim.

Reliability comes first

Reliability is consistency. If the same person sat the same test twice under the same conditions, how close would the two scores be? No test is perfectly consistent, and the useful question is how much of a score is signal and how much is noise.

Two measures matter in practice. Internal consistency asks whether items within a test hang together — whether they are measuring the same thing. Test-retest reliability asks whether scores hold up over time. Both are usually reported as a coefficient between 0 and 1, and for decisions about individuals rather than groups, anything below about 0.80 deserves questions.

More useful than the coefficient itself is the standard error of measurement, which converts reliability into the units you actually use. A test with a standard error of four percentile points means a candidate scoring at the 60th percentile might reasonably have scored anywhere in the mid-50s to mid-60s. That should change how you treat small differences between candidates — because most of the time, they are not differences at all.

Validity is about the inference, not the test

The common phrasing — "this is a valid test" — is technically wrong, and the error matters. Validity is a property of the inference you draw from a score, in a particular context, for a particular purpose. A test can support one inference well and another badly.

Criterion validity asks whether scores predict something you care about, usually job performance. Construct validity asks whether the test measures the thing it claims to. Content validity asks whether items adequately sample the domain.

One important recent shift: the validity estimates quoted across the industry for decades were revised downward in 2022, after Sackett and colleagues showed that standard corrections for range restriction had been systematically overcorrecting. General mental ability remains a strong predictor, but the headline figures many suppliers still quote are inflated. A supplier citing pre-2022 numbers without acknowledging the revision is either behind or hoping you are.

Fairness is measurable, not asserted

Fairness has a technical meaning here, and it is not the same as everyone scoring equally. It means the test predicts equally well for different groups, and that items do not behave differently for candidates of equal underlying ability.

Differential item functioning analysis identifies individual items that behave oddly for particular groups — an item where two candidates of the same ability but different backgrounds have different chances of answering correctly. Adverse impact analysis looks at outcomes: whether selection rates differ across groups, and if so, whether the test is justified for the role.

Group differences in scores are not automatically unfair. A test can show differences and still predict performance equally well for everyone. But those differences create a legal obligation to demonstrate the test is a proportionate means of achieving a legitimate aim — which requires evidence, not confidence.

What to ask a supplier

Ask for the technical manual, not the brochure. It should report reliability with the sample it was calculated on, validity evidence with the criterion and the study design, and fairness analyses by group.

Ask when the norms were collected and on whom. Ask whether the test has been independently reviewed — in the UK, the British Psychological Society operates a registration and review scheme, and published reviews are exactly the independent scrutiny that marketing copy is not.

And ask what the standard error is. A supplier who cannot tell you how much noise is in their scores has not thought hard enough about what those scores mean.

The short version
  • Reliability is consistency — for decisions about individuals, below about 0.80 deserves questions
  • Validity belongs to the inference you draw, not to the test itself
  • Industry validity figures were revised downward in 2022; check the date on any quoted to you
  • Ask for the technical manual, the norm collection date, and the standard error
Get new articles by email

One email a month on assessment best practice. No spam, unsubscribe anytime.

© 2026 iTestHub