Situational judgement tests present candidates with realistic work dilemmas and ask them to judge the possible responses. They sit between the abstraction of an ability test and the expense of an assessment centre — and the research explains why they have become a fixture of high-volume and high-stakes selection alike.
What SJTs predict
Meta-analyses put the criterion-related validity of SJTs against job performance at around .26 (with some estimates for specific designs higher).[1] More important than the headline figure is the incremental evidence: SJTs add 3–5% predictive value over cognitive ability, 6–7% over personality, and 1–2% over both combined[2] — meaning they capture something the other instruments do not. The literature identifies that something as practical, interpersonal judgement: SJTs measure how candidates weigh competing interests, values and consequences in complex social situations, a domain that ability tests reach only obliquely and self-report questionnaires not at all.[3]
SJTs are not construct-pure: they draw on cognitive ability, conscientiousness, agreeableness and emotional stability at once.[1] Researchers increasingly describe them as measures of general domain knowledge — knowing what effective behaviour looks like in a class of situations.[4] For hiring, that blend is a feature: the test samples the judgement the job will actually demand.
Design decisions the evidence settles
Response instructions. Asking what one should do (knowledge instructions) versus what one would do (behavioural tendency) changes what the test measures: knowledge formats correlate more with cognitive ability, tendency formats more with personality.[1] Knowledge instructions are also more resistant to faking, which matters in applicant settings; a large high-stakes study found response instructions did not meaningfully change criterion validity.[5]
Fidelity. Presentation affects validity: converting a video-based SJT to text, holding content constant, significantly reduced its criterion validity and made it more of a reading test.[6] Richer, more realistic presentation earns its production cost.
Scoring. Scoring keys built from expert consensus, and scoring methods chosen deliberately, measurably affect validity[7] — an SJT is only as good as the judgement standard behind it, which is why generic off-the-shelf scenarios underperform tests built with subject-matter experts for the actual role.
Why candidates and organisations tolerate them well
Across the literature, SJTs show a rare combination of practical advantages: favourable applicant reactions (the scenarios visibly resemble the job), comparatively low susceptibility to faking, and — in several studies — smaller subgroup differences than cognitive ability tests, easing the validity–diversity trade-off.[8] This is why SJTs carry heavy weight in medical and public-sector selection, where fairness scrutiny is highest.[9]
The fairness evidence is strongest for neurodivergent and disabled candidates. A UK evaluation of 452,933 candidates who completed an SJT in live recruitment found little to no differences in performance across disability and neurodivergence categories — including dyslexia, autism, ADHD, dyspraxia and mental health conditions — suggesting well-designed SJTs are unlikely to create disparities in outcomes for these groups.[10] An earlier study reached the same conclusion.[10] The caveat cuts the other way: format still matters. A UK Employment Appeal Tribunal upheld a finding of indirect disability discrimination against a government department's multiple-choice test where a candidate with an autistic spectrum condition was refused a short-written-answer adjustment[11] — the fairness advantage belongs to SJTs that are thoughtfully designed and offer reasonable adjustments, not to the format automatically.
The caveats: SJT scores are harder to interpret than a single-construct test precisely because they blend constructs; retest reliability varies with design; and a scenario set anchored to one organisation's values may not transfer to another. These argue for careful, role-specific development rather than against the method.
Where SJTs fit in a selection battery
The evidence supports using SJTs as the judgement layer of a battery: ability tests for learning potential, skills tests for current capability, personality for behavioural tendencies, and an SJT for whether the candidate recognises effective behaviour in the situations the role will produce. Because its incremental validity over the other three is modest,[2] an SJT earns its place mostly where judgement is the job — customer-facing, safety-critical, supervisory and values-laden roles — and where its candidate-experience and fairness advantages carry weight at scale.
- 1. McDaniel, M., Hartman, N., Whetzel, D. & Grubb, W. (2007). Situational judgment tests, response instructions, and validity: a meta-analysis. Personnel Psychology 60.
- 2. McDaniel, M. et al. (2007), incremental validity analyses, summarised in Lievens, Peeters & Schollaert (2008). Situational judgment tests: a review of recent research. Personnel Review 37.
- 3. Christian, M., Edwards, B. & Bradley, J. (2010). Situational judgment tests: constructs assessed and a meta-analysis of their criterion-related validities. Personnel Psychology 63.
- 4. Motowidlo, S. et al.; Lievens, F. & Motowidlo, S. (2016). Situational judgment tests: from measures of situational judgment to measures of general domain knowledge. Industrial and Organizational Psychology 9.
- 5. Lievens, F., Sackett, P. & Buyse, T. (2009). The effects of response instructions on situational judgment test performance and validity in a high-stakes context. Journal of Applied Psychology 94.
- 6. Lievens, F. & Sackett, P. (2006). Video-based versus written situational judgment tests. Journal of Applied Psychology 91.
- 7. Weng, Q., Yang, H., Lievens, F. & McDaniel, M. (2018). Optimizing the validity of situational judgment tests: the importance of scoring methods. Journal of Vocational Behavior 104.
- 8. Whetzel, D. et al. (2008); Bauer, T. & Truxillo, D. (2006); Kasten, N. et al. (2020) — subgroup differences, applicant reactions, and faking resistance; summarised in Kepes, Keener, Lievens & McDaniel (2025), Journal of Management.
- 9. Patterson, F. et al. (2016); Webster, E. et al. (2020). Situational judgement test validity for selection: systematic review and meta-analysis. Medical Education 54.
- 10. Shalfrooshan, A. et al. (2023). Neurodiversity, disability and assessment: evaluation of the impact of SJTs during recruitment (N = 452,933); Willis, C. et al. (2021), similar findings of no significant performance differences.
- 11. Government Legal Service v Brookes (2017), UK Employment Appeal Tribunal.
We design situational judgement tests with your subject-matter experts, scored against an expert consensus key. Get in touch to discuss.
Enquire about assessments