Pre-employment testing has a strong evidence base, and most of the useful findings are decades old rather than new. The core question is simple: which assessments actually predict how well someone will do the job, and which just feel rigorous? This guide works through the research, ranks the main test types by what they predict, and sets out how to use them legally in both the US and the UK. It also explains, honestly, where an applicant tracking system like Treegarden helps and where it does not.
Why Structured Assessment Beats Gut Feel
The case for testing rests on a comparison with the default alternative: the unstructured interview. Decades of evidence show that a freewheeling chat is one of the weaker ways to predict performance, because it is vulnerable to first impressions, similarity bias and the halo effect, where one strong impression colours the whole judgement. Structured, job-relevant assessment reduces that noise by asking every candidate to demonstrate the same job-related skills against the same criteria.
This is not a claim about a particular vendor's software; it is a claim about method. A structured scoring rubric, a work sample that mirrors real tasks, or a validated cognitive test all share the same advantage: they make candidates comparable on something that matters for the role, rather than on who interviewed more smoothly. The rest of this guide is about choosing the right method for the right job and applying it lawfully.
Types of Pre-Employment Tests: What the Research Says
The most cited source on this is Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin, which summarised 85 years of research on selection methods. It reported that general mental ability (cognitive ability) predicted job performance at roughly r = .51, that work sample tests and structured interviews were also among the strongest predictors, and that unstructured interviews trailed at about .38. An equally weighted combination of a cognitive test and a structured interview produced one of the highest validities of all, around .63. For two decades these figures anchored best practice.
Important: the numbers were revised in 2022
Sackett, Zhang, Berry and Lievens (2022) showed that earlier meta-analyses had over-corrected for range restriction, inflating some validities, particularly for cognitive ability. Their revised estimates place structured interviews highest among common methods at about r = .42, with job knowledge tests near .40, work samples near .33 and cognitive ability near .31. The ranking shifted: job-specific, structured methods now look at least as predictive as raw cognitive testing. Cite the revised figures, not the 1998 originals, when you need current numbers.
Read together, the two bodies of work give a defensible hierarchy for most roles: structured interviews, work samples and job knowledge tests are the dependable core; cognitive ability remains a useful predictor but a more modest one than once believed; and personality measures are supporting evidence, not the main event. The practical takeaway is to match the method to the job. A software developer is better assessed with a realistic coding work sample than a personality questionnaire; a customer-service role might pair a short structured interview with a job-relevant scenario.
Cognitive Ability Tests: Strong, but Not a Silver Bullet
Cognitive ability tests measure reasoning, problem-solving and the capacity to learn quickly. They predict performance across a wide range of roles, which is why they were long treated as the single best predictor. The 2022 re-analysis tempers that: cognitive ability is still a meaningful predictor, but its corrected validity is lower than the 1998 estimate, and structured, job-specific methods match or exceed it.
Cognitive tests also carry the clearest legal risk. They tend to show larger score differences between demographic groups than work samples or structured interviews, which raises adverse-impact exposure under the US Uniform Guidelines and the UK Equality Act 2010. Notably, the Sackett team found that pairing or replacing cognitive tests with job-specific methods can preserve predictive power while reducing that adverse impact. If you use cognitive testing, keep it job-relevant, combine it with other evidence, and monitor outcomes by group.
Skills and Work Sample Tests: Assessing What the Job Requires
Work sample and skills tests ask candidates to perform tasks that resemble the actual work. They are consistently among the strongest predictors precisely because they measure job-relevant behaviour directly, and candidates tend to see them as fair, which helps completion rates. Examples by role family:
- Technical roles: a scoped coding exercise or a realistic debugging task.
- Creative roles: a short, time-boxed brief that mirrors real deliverables.
- Customer-facing roles: a written response or role-play scenario scored against a rubric.
Treegarden's role here is supporting, and worth stating plainly. Treegarden does not generate or auto-grade skills tests, and it does not recommend which test to run or how long it should be. What it does provide is the surrounding workflow: bulk CV upload and AI candidate matching to shortlist applicants faster, structured interviews with consistent scoring, and a Kanban pipeline where you can record assessment outcomes and move candidates through stages. The test itself you design or source; Treegarden helps you run a consistent, organised process around it.
Personality Tests: What They Can and Cannot Tell You
Personality assessments based on the Big Five (the OCEAN model) can add incremental information, but they are weaker predictors than cognitive ability, work samples or structured interviews and should never be the deciding factor. Conscientiousness is the trait most consistently associated with job performance across roles; other traits matter more situationally, such as agreeableness in team-heavy work.
- Reasonable use: as one supporting input, especially for traits with a clear job rationale.
- Poor use: as a primary gate, or for safety-critical decisions that demand demonstrated skill.
- Watch for faking: self-report measures can be gamed, so corroborate with behaviour-based evidence.
A practical rule is to pair any personality measure with a work sample or structured interview, and to favour validated instruments over popular type-based questionnaires that lack predictive-validity evidence. Treat personality data as a prompt for better questions, not as a verdict.
Legal Considerations: Keeping Assessments Defensible
Any test used in hiring is a selection procedure, and it must be job-related and consistent with business necessity. The compliance essentials differ by jurisdiction:
- US: the EEOC's Uniform Guidelines on Employee Selection Procedures, the four-fifths (80%) rule for adverse impact, ADA Title I reasonable accommodations, and OFCCP obligations for federal contractors.
- UK: the Equality Act 2010, which prohibits both direct and indirect discrimination, plus lawful handling of candidate data under UK GDPR.
The single most important habit is adverse-impact monitoring. Validate that each test predicts performance for the role, document that rationale, and check selection rates across groups.
The four-fifths rule, in practice
Under the US Uniform Guidelines, if the selection rate for any protected group is less than four-fifths (80%) of the rate for the highest-scoring group, that is treated as evidence of adverse impact. For example, if 50% of one group passes a test but only 30% of another does, the ratio is 0.6, below the 0.8 threshold, and the test needs justification or revision. UK employers face a parallel duty to avoid indirect discrimination under the Equality Act 2010.
How Testing Fits Into Your ATS Pipeline
Whatever assessments you choose, they have to live inside a hiring workflow or they create chaos: results in spreadsheets, candidates lost between stages, inconsistent scoring. This is where an ATS earns its place, and where Treegarden's contribution is concrete:
What Treegarden does in the testing workflow
Shortlist with bulk CV upload and AI candidate matching (Edera AI), run structured interviews with consistent scoring, and track each candidate on a Kanban pipeline where assessment outcomes and interview notes sit in one place. Treegarden does not author, deliver or auto-score the tests themselves; you bring or build those, and Treegarden keeps the process organised and consistent.
Treegarden is built for UK and US small and mid-sized businesses, with pricing from $299 per month (Startup), $499 per month (Growth) and $899 per month (Scale) in USD, or £235, £395 and £710 per month in the UK, and custom Enterprise pricing. There is no free plan; you can evaluate it through a guided demo with a sandbox on request. Used well, the ATS removes the administrative friction around assessment so your hiring team can focus on judging the evidence the tests produce.
Scoring and Benchmarking Pre-Employment Test Results
Collecting test data is only half the equation - interpreting it correctly is where organisations most often go wrong. Without benchmarks, a raw score of 72% on a cognitive test tells you nothing meaningful. Is that strong, average, or weak relative to the population of people likely to apply for this role? The answer shapes your hiring decision entirely.
There are three benchmarking approaches used in practice. The first is normative benchmarking, where each candidate's score is compared against a reference population - typically a large sample of people in similar roles or at similar career stages. Most commercial test providers publish norm tables segmented by job family, seniority level, and industry. A score in the 70th percentile means the candidate outperformed 70% of the reference population. Normative benchmarks are useful when your primary goal is ranking candidates relative to each other.
The second approach is criterion-referenced scoring, where you set a minimum pass threshold based on what the job actually requires. If the role demands mid-level Excel proficiency, you define what "mid-level" looks like and score candidates pass/fail against that standard. This works well for skills tests with clear, objectively defined competency levels and avoids the ranking competition dynamic that can disadvantage groups who are broadly qualified but score slightly below average on a normative scale.
The third approach - less common but most rigorous - is predictive validity scoring. Here, you compare test scores from past hires against their subsequent job performance ratings to determine which score ranges actually predict success in your specific context. Organisations that have been testing candidates for two or more years and have strong performance data can build these internal benchmarks. They are far more accurate than generic industry norms because they reflect the actual demands of your culture and operating environment.
Score weighting also matters. A common mistake is treating a personality assessment with the same weight as a cognitive or skills test, despite the latter having far stronger predictive validity evidence. A sensible weighting framework might assign cognitive ability 30%, job-specific skills tests 40%, structured interview 20%, and personality assessment 10%. Documenting your weighting rationale protects you legally and makes the process transparent to candidates who request feedback.
Candidate Experience and Test Completion Rates
Pre-employment tests create friction in the application process. Every additional step can reduce your candidate completion rate, sometimes significantly. Assessments added early in the application process tend to depress completion, with the drop most severe for in-demand candidates who have multiple offers in progress and little incentive to invest time in any single application. The research is more nuanced than "shorter is always better": a HireVue analysis of more than 30 million pre-hire assessments found that overall length is a weak predictor of completion, and that most dropout happens in the first five to ten minutes, which means how you frame and open the assessment matters more than trimming a few minutes off the end.
Managing this friction requires deliberate design. Timing matters: placing tests after an initial screening conversation, rather than immediately after application submission, dramatically improves completion rates. Candidates who have already invested time in a phone screen are far more committed and less likely to abandon. For high-volume roles where individual screening is impractical, keep the early-stage test short - under 20 minutes for the initial gate.
Transparency also drives completion. Candidates who understand why they're being tested, how long it will take, how their data will be used, and how results feed into the decision are significantly more likely to complete the assessment and report a positive experience. A brief pre-test briefing email that explains the process, sets time expectations, and reassures candidates that the test is one input among many - not a pass/fail gate - reduces anxiety and improves engagement.
Mobile optimisation is increasingly critical. In competitive talent markets, a meaningful proportion of candidates apply and complete assessments on mobile devices. Tests that are not designed for mobile have lower completion rates and introduce a demographic bias, as mobile-first applicants skew younger and toward certain socioeconomic groups. Before deploying any assessment tool, verify that the candidate-facing interface works correctly on smartphones.
Finally, provide feedback where possible. Candidates who receive score feedback - even a brief summary - report significantly higher satisfaction with the hiring process regardless of the outcome. This has direct employer branding value: candidates who are treated respectfully through an unsuccessful application are more likely to apply again, refer others, and speak positively about the organisation. In tight talent markets, that reputational effect compounds over time.
Free Calculators for This Topic
Save time with these free HR calculators - no sign-up required:
Frequently Asked Questions
Why are work samples and structured methods better than unstructured interviews?
They directly assess job-relevant behaviour, which reduces the halo effect that inflates a confident but unqualified candidate in a freewheeling chat. In Schmidt and Hunter's 1998 meta-analysis, structured interviews predicted performance at r = .51 versus .38 for unstructured interviews. The 2022 Sackett et al. re-analysis, which corrected a range-restriction error, placed structured interviews highest among common methods at about r = .42.
How do I keep pre-employment tests legally defensible?
Use job-related, validated tests, document why each test maps to the role, and run adverse-impact analysis against the four-fifths rule in the US Uniform Guidelines. US employers should plan for ADA reasonable accommodations and keep records for EEOC or OFCCP review; UK employers must avoid indirect discrimination under the Equality Act 2010 and handle candidate data lawfully under UK GDPR.
Can personality tests be used for hiring?
Yes, if they are job-relevant and validated, but they are weaker predictors than cognitive ability, work samples or structured interviews and should support rather than drive a decision. Conscientiousness is the Big Five trait most consistently linked to performance. Avoid unvalidated type-based questionnaires that lack predictive-validity evidence.
How long should pre-employment tests take?
There is no official standard, but in practice keeping early-stage assessments short, often under 20 to 30 minutes, protects completion rates, while more senior roles can justify longer work samples. Match the length to what the role genuinely requires, and place longer tests after an initial screening conversation to reduce drop-off.
Used well, pre-employment testing moves hiring from impression to evidence, and the research points to a clear core: structured interviews, work samples and job knowledge tests, with cognitive ability as a useful supporting predictor and personality measures as a minor input. Treegarden does not author or grade the tests themselves; it gives UK and US small and mid-sized businesses the workflow around them, including bulk CV upload, AI candidate matching, structured interviews and a Kanban pipeline, with transparent pricing from $299 per month and no free plan. To see how that process works against your own roles, book a demo.