multi-item-scale · reviewed 2026-07-12 contested: Structural validity, Measurement invariance

Kessler Psychological Distress Scale (K10)

Licence verified: 2026-07-12 · record reviewed: 2026-07-12 (pass two)

Identity

Version: K10 (10-item); a 6-item short form (K6) is embedded within it. Both derive from the same item-response-theory development work.

Structure: 10 items, each rated on a five-point frequency scale over the past 30 days; total range 10 to 50 (or 0 to 40 under the alternative 0 to 4 coding).

Original citation: Kessler RC, Andrews G, Colpe LJ, et al. Short screening scales to monitor population prevalences and trends in non-specific psychological distress. Psychological Medicine 2002;32(6):959-976. doi:10.1017/s0033291702006074

Steward / publisher: Ronald C. Kessler and colleagues, Department of Health Care Policy, Harvard Medical School. Distributed through the National Comorbidity Survey site (hcp.med.harvard.edu/ncs/k6_scales.php). The K10 itself was developed with the Clinical Research Unit for Anxiety and Depression (CRUFAD), Australia.

Licence status (verified 2026-07-12): Free to use with no formal permission or registration required. The steward's current distribution page (Harvard Medical School National Comorbidity Survey, 'Kessler Psychological Distress Scale (K10)') states under its permission-requests heading that use of the K6 and K10 is free and does not require any formal permission or approval, asking only that users cite the source article and include the copyright notice when using the scales. Copyright is held by Ronald C. Kessler. No named open-content licence identifier (for example a Creative Commons code) is asserted by the steward; the scale is distributed as a free-to-use, cite-and-attribute instrument rather than under a formal open licence. Redistribution and modification terms are not explicitly addressed on the page, so verify with the steward before either.
Source: Harvard Medical School, National Comorbidity Survey, 'Kessler Psychological Distress Scale (K10)' distribution page, https://www.hcp.med.harvard.edu/ncs/k6_scales.php. The page body was read in the latest review pass; its permission-requests section states that use of the K6 and K10 is free and requires no formal permission or approval, asking users to cite the article and include the copyright. Not sourced from any founding paper or review.

Constructs claimed

Non-specific psychological distress over the preceding 30 days, indexed through symptoms of anxiety and depression (nervousness, agitation, fatigue, hopelessness, negative affect). It is a dimensional screen for the likely presence of a common mental disorder, not a diagnostic instrument and not a measure of any single named disorder.

Evidence

Deployment context caveat. The K10 is an epidemiological and clinical screening instrument developed for population surveys and validated against structured psychiatric diagnosis. It is not a workplace-designed measure and carries no evidence base against workplace outcomes. Any organisational deployment is off-label relative to its validation: scores index general psychological distress, not work-related distress, and no occupational cut-offs or work-outcome linkages have been established. 'Clinical' or 'diagnostic' language is appropriate only to its origin and reference-standard validation, never to a workplace-screening use. (applies to every property below)

Structural validity Moderatecontestedevidence form: canonical

indirect . Structural evidence is abundant but earned mainly in Australian, sub-Saharan African, Chinese and South American general or clinical samples; no UK working-adult factor-analytic study was located in the latest review pass.

The dimensionality of the K10 is genuinely contested. The developers treated it as essentially unidimensional (a single distress continuum) from item-response-theory modelling (Kessler 2002). A widely cited community-sample analysis instead found four first-order factors (labelled nervous, negative affect, fatigue and agitation) resolving into two second-order factors interpreted as depression and anxiety, replicated across two survey waves (Brooks 2006). Subsequent confirmatory work has not settled on one model: a large Chinese healthcare-professional sample favoured a two-factor (depression and anxiety) oblique model over a one-factor model (Wang 2025), a two-factor solution also emerged among children of Chinese migrant workers (Ren 2021) and in a Portuguese adult sample the original one-dimensional structure was not confirmed in favour of correlated anxiety and depression factors (Pereira 2019). Several sub-Saharan validations converged instead on a unidimensional model with correlated error terms as the best-fitting and most parsimonious solution, with four-factor solutions judged to be overfitted or artefactual (Milkias 2022; Naisanga 2022; Hoffman 2022). Bifactor models (a general distress factor plus specific anxiety and depression factors) have fitted well in Brazilian samples (Perrelli 2024; Peixoto 2021). Critically, neither the single-factor nor the multifactor model fitted in a treatment-seeking clinical sample, which the authors read as a warning against assuming the epidemiological structure holds in clinical settings (Berle 2010). The pattern is best summarised as a robust general distress dimension overlaid by a separable anxiety and depression split, with the winning statistical model depending heavily on sample, estimator and whether correlated errors are permitted.

Convergent and discriminant validity Moderatewell-establishedevidence form: canonical

indirect . Convergent evidence comes from Australian, Canadian, Brazilian and other non-UK general, student and adolescent samples rather than UK workers.

The K10 correlates strongly with other distress and internalising measures. It showed a strong correlation with the Self-Reporting Questionnaire (r = 0.81) in Brazilian higher-education students (Perrelli 2024) and a moderate correlation (r = 0.63) with the emotional-symptoms subscale of the Strengths and Difficulties Questionnaire in Australian adolescents (Blake 2023). Latent-variable work places the K10 firmly on a single internalising-psychopathology dimension shared with DSM-IV depression and anxiety diagnoses (Sunderland 2013). Higher K10 scores track strongly and monotonically with suicidal ideation, those in the very-high band being an order of magnitude more likely to report ideation than those in the low band (Chamberlain/Goldney 2009). Discriminant separation from unrelated constructs is less formally documented; most validation studies emphasise convergent rather than discriminant coefficients.

Criterion validity: reference standard Highwell-establishedevidence form: canonical

indirect . The strongest diagnostic-accuracy evidence is Australian, US and Canadian general-population and military samples; no UK working-adult criterion study against a diagnostic standard was located, and accuracy varies by setting.

This is the K10's strongest evidence base. Against structured diagnostic interviews (CIDI or SCID for DSM-IV disorders) the scale discriminates cases from non-cases well. In the original development and clinical reappraisal, areas under the ROC curve were 0.87 to 0.88 for disorders of moderate severity and 0.95 to 0.96 for severe disorders (Kessler 2002). In the nationally representative Australian National Survey of Mental Health and Well-Being the K10 achieved an AUC of 0.90 (95% CI 0.89 to 0.91) for CIDI/DSM-IV mood and anxiety disorders, outperforming the GHQ-12 (AUC 0.80) (Furukawa 2003). Against serious mental illness defined by SCID plus functional impairment, the K10 AUC was 0.85 (Kessler 2003). Performance is more modest outside high-income settings: in the South African Stress and Health study the K10 showed only moderate discrimination (AUC 0.73 for depression, 0.72 for anxiety) and failed the authors' joint sensitivity and positive-predictive-value criteria, with poorer discrimination in the Black subgroup (Andersen 2011). In Canadian military personnel the AUC against four past-month disorders was 0.92, with a screening cut-off of 10 or greater giving 86% sensitivity and 83% specificity (Sampasa-Kanyinga 2018). A Swiss community study found much poorer agreement with MINI diagnoses than the classic Australian work, with low Cohen's kappa at every cut-off, cautioning against community screening use (Osman 2022). Cut-offs are population-dependent and not universal, and the reference-standard time frame (12-month diagnosis) often mismatches the K10's 30-day window, a recognised limitation (Andersen 2011).

Criterion validity: organisational Absent (a finding about the literature)untestedevidence form: canonical

indirect . Even the associational workplace literature is largely Australian, US and Japanese; no UK workplace criterion evidence was located and none is criterion-validation in design.

No study located in the latest review pass validated the K10 against an organisational or work outcome (sickness absence, turnover, job performance, or work-recorded diagnosis) as a criterion. The K10 and its K6 short form appear widely in occupational and workplace surveys, but as an exposure or prevalence measure rather than as a screen validated against a work-outcome standard. Examples of such descriptive use include an Australian industry comparison linking very-high distress to self-reported productivity loss and work-cutback days (Burns 2023), a thirty-seven-year US panel study relating occupation and job tenure to new distress cases measured with the K6 (Laditka 2023), and a Japanese occupational cohort using a K6 cut-off to define distress caseness at baseline ([workplace social support, cited in record notes]). These establish that K10/K6 distress associates with work-relevant variables, but none provides criterion validity in the COSMIN sense: distress is the predictor or outcome of interest, not a test being validated against an independent work-outcome gold standard, and no work-based cut-off has been derived or calibrated. Organisational criterion validity is therefore Absent as a formal psychometric property.

Internal consistency Highwell-establishedevidence form: mixed

indirect . Alpha is high almost everywhere it has been measured, but the pooled evidence is dominated by non-UK general, clinical, student and military samples; no UK working-adult alpha was isolated in the latest review pass, though a UK adult study using the K6 short form reported adequate fit ([Lantos 2023](https://doi.org/10.1016/j.jad.2023.06.033)).

Internal consistency is consistently high and this is the K10's best-replicated reliability property. A reliability-generalisation meta-analysis of 48 studies (2002 to 2024) estimated a pooled Cronbach's alpha of 0.90 (95% CI 0.88 to 0.91) for the K10, with variation across populations from about 0.78 (a Tanzanian sample) to 0.97 (Australian samples), and highest values among adolescents (0.93) and carers (0.91) (Wojujutari 2024). Primary studies corroborate this: alpha 0.88 in Canadian military personnel (Sampasa-Kanyinga 2018), 0.95 in Chinese healthcare professionals (Wang 2025), 0.83 to 0.86 across Ethiopian, Ugandan and South African general and medical samples (Milkias 2022; Naisanga 2022; Hoffman 2022, the last also reporting McDonald's omega total of 0.88), and 0.91 in Portuguese adults (Pereira 2019). Lower values appear in some community samples, for example alpha 0.81 in a Swiss young-adult sample (Osman 2022). Values in the 0.80s to low 0.90s are the norm.

Test-retest reliability Absent (a finding about the literature)thin

indirect . No canonical K10 test-retest coefficient was located in the latest review pass; the absence itself is population-general.

No test-retest or temporal-stability coefficient for the canonical K10 (or K6) was located in in the latest review pass's searches. The K10 was developed as a cross-sectional epidemiological screen with a 30-day recall window, and its reliability evidence base is overwhelmingly internal-consistency (alpha) based; the large reliability-generalisation meta-analysis synthesised alpha only, not retest reliability (Wojujutari 2024). This is a genuine and important gap: high internal consistency does not establish temporal stability, and the two must not be conflated. For an instrument used to track change (including any longitudinal workplace use), the absence of established test-retest reliability is a material limitation and is recorded here as the finding it is, not averaged away.

Measurement invariance Moderatecontestedevidence form: mixed

indirect . Sex and age invariance is reasonably supported but in Australian, Chinese and Brazilian samples; cross-cultural equivalence is contested and occupational invariance is untested.

Invariance evidence is moderate and mostly supportive across sex and age, thinner across culture and untested across occupation. Strict measurement invariance held across two Australian survey administrations ten years apart and across age bands for a latent internalising model incorporating K10 items (Sunderland 2013). Full measurement invariance across gender was supported among children of Chinese migrant workers (Ren 2021), and Rasch analysis in older Australians found scale invariance across sex, age and education after minor model modification (Calkin 2023). A Brazilian adaptation reported multiple-group invariance by gender and age range (Peixoto 2021). A UK adult study reported support for measurement invariance across birth sex and age for the K6 short form (Lantos 2023). Against this, cross-national and cross-ethnic comparability is questioned: the South African study found significantly lower discrimination in the Black subgroup, implying non-equivalence by race/ethnicity (Andersen 2011). No study located in the latest review pass tested invariance across occupational groups or between working and non-working adults, which is the comparison most relevant to workplace deployment.

Sub-grades (evidence differs by subgroup):

  • {"subgroup": "across sex", "grade": "Moderate", "note": "Full or strict invariance supported in several samples (Ren 2021; Calkin 2023; Sunderland 2013)."}
  • {"subgroup": "across age", "grade": "Moderate", "note": "Invariance supported across age bands and over a ten-year interval in Australian data (Sunderland 2013; Calkin 2023)."}
  • {"subgroup": "across culture/ethnicity", "grade": "Low", "note": "Differential performance by race/ethnicity in South Africa suggests non-equivalence (Andersen 2011)."}
  • {"subgroup": "across occupation", "grade": "Absent", "note": "No test of invariance across occupational groups or working vs non-working adults located in the latest review pass."}

Responsiveness and MIC Absent (a finding about the literature)thinevidence form: canonical

indirect . No responsiveness or MIC evidence located in any population, UK or otherwise.

No formal responsiveness study (ability to detect within-person change) or minimal important change (MIC) value for the K10 was located in the latest review pass. Related psychometric work is severity-oriented rather than change-oriented: severity banding into no, mild, moderate and severe distress is well established for cross-sectional interpretation (Ul Husnain 2024, using Australian HILDA data), and Rasch work has produced ordinal-to-interval conversion tables intended to improve measurement precision in older adults (Calkin 2023), but neither reports a minimal important change or an anchor-based responsiveness estimate. The K10's design and evidence base are cross-sectional; responsiveness and MIC are effectively unestablished.

Populations, languages and norms

Norms and translations are extensive. Australian normative data are the reference standard: interpretive norms from the 1997 survey (Andrews 2001) and full normative tables by sex, age and disorder status from the 2007 National Survey of Mental Health and Wellbeing (n = 8841), with stratum-specific likelihood ratios for estimating disorder probability (Slade 2011). Australian adolescent norms (ages 11 to 17, n = 2964) are available with the caveat of low specificity for major depressive disorder (Blake 2023). Multi-country European general-population norms across seven countries (n = 7087) show mean total scores varying from about 6.9 (Netherlands) to 9.9 (Spain) on the 0 to 40 metric, women and younger adults scoring higher (Lehmann 2023). Validated translations located in the latest review pass include Arabic ([Easton 2017, cited in record notes]), Brazilian Portuguese (Perrelli 2024; Peixoto 2021), European Portuguese (Pereira 2019), and multiple sub-Saharan African adaptations (Milkias 2022; Naisanga 2022; Hoffman 2022). UK usage is real: the K10 and K6 appear in UK survey and social-science research, and a UK adult sample analysed the K6 short form (Lantos 2023); however, no UK working-adult normative table for the full K10 was located in the latest review pass. For a UK working-adult audience, the closest directly usable norms are the seven-country European set, which does not include the UK.

Criticisms and controversies

Three issues recur. First, dimensionality is unresolved: the developers' unidimensional treatment coexists with well-supported two-factor (anxiety and depression) and bifactor models, and the best-fitting model is sample- and method-dependent (Brooks 2006; Wang 2025; Milkias 2022). Second, the epidemiological factor structure and cut-offs do not transfer cleanly to clinical or community-screening settings: neither standard model fitted a treatment-seeking sample (Berle 2010), and a Swiss community study found agreement with diagnostic interviews too poor to recommend community screening (Osman 2022). Third, cut-offs are not universal, vary by country, age and sex, and diagnostic accuracy drops in some non-Western and minority-ethnic samples (Andersen 2011; Blake 2023). Under-recognised gaps are the near-absence of published test-retest reliability and of responsiveness or minimal-important-change evidence, both material for any use that tracks change over time. For workplace use specifically, the instrument has no work-outcome validation and no occupational norms or invariance testing.

References (30)

  1. Kessler R C, Andrews G, Colpe L J et al. (2002). Short screening scales to monitor population prevalences and trends in non-specific psychological distress. https://doi.org/10.1017/s0033291702006074
  2. Andrews G, Slade T (2001). Interpreting scores on the Kessler Psychological Distress Scale (K10). https://doi.org/10.1111/j.1467-842x.2001.tb00310.x
  3. Kessler Ronald C, Barker Peggy R, Colpe Lisa J et al. (2003). Screening for serious mental illness in the general population. https://doi.org/10.1001/archpsyc.60.2.184
  4. Furukawa T A, Kessler R C, Slade T et al. (2003). The performance of the K6 and K10 screening scales for psychological distress in the Australian National Survey of Mental Health and Well-Being. https://doi.org/10.1017/s0033291702006700
  5. Brooks Robert T, Beard John, Steel Zachary (2006). Factor structure and interpretation of the K10. https://doi.org/10.1037/1040-3590.18.1.62
  6. Berle David, Starcevic Vladan, Milicevic Denise et al. (2010). The factor structure of the Kessler-10 questionnaire in a treatment-seeking sample. https://doi.org/10.1097/NMD.0b013e3181ef1f16
  7. Andersen L S, Grimsrud A, Myer L et al. (2011). The psychometric properties of the K10 and K6 scales in screening for mood and anxiety disorders in the South African Stress and Health study. https://doi.org/10.1002/mpr.351
  8. Slade Tim, Grove Rachel, Burgess Philip (2011). Kessler Psychological Distress Scale: normative data from the 2007 Australian National Survey of Mental Health and Wellbeing. https://doi.org/10.3109/00048674.2010.543653
  9. Sunderland Matthew, Slade Tim, Carragher Natacha et al. (2013). Age-related differences in internalizing psychopathology amongst the Australian general population. https://doi.org/10.1037/a0034562
  10. Ren Qiang, Li Yong, Chen Ding-Geng (2021). Measurement invariance of the Kessler Psychological Distress Scale (K10) among children of Chinese rural-to-urban migrant workers. https://doi.org/10.1002/brb3.2417
  11. Lehmann J, Pilz M J, Holzner B et al. (2023). General population normative data from seven European countries for the K10 and K6 scales for psychological distress. https://doi.org/10.1038/s41598-023-45124-0
  12. Blake Julie A, Farugia Taya L, Andrew Brooke et al. (2023). The Kessler Psychological Distress Scale in Australian adolescents: Analysis of the second Australian Child and Adolescent Survey of Mental Health and Wellbeing. https://doi.org/10.1177/00048674231216601
  13. Calkin Cailen J, Numbers Katya, Brodaty Henry et al. (2023). Measuring distress in older population: Rasch analysis of the Kessler Psychological Distress Scale. https://doi.org/10.1016/j.jad.2023.02.116
  14. Wang Ye, Zeng Zheng, Huang Changqun et al. (2025). Large-scale validation of the Kessler-10 Scale's psychometric properties among healthcare professionals in China. https://doi.org/10.1016/j.genhosppsych.2025.02.017
  15. Wojujutari Ajele Kenni, Idemudia Erhabor Sunday (2024). Consistency as the Currency in Psychological Measures: A Reliability Generalization Meta-Analysis of Kessler Psychological Distress Scale (K-10 and K-6). https://doi.org/10.1155/2024/3801950
  16. Sampasa-Kanyinga Hugues, Zamorski Mark A, Colman Ian (2018). The psychometric properties of the 10-item Kessler Psychological Distress Scale (K10) in Canadian military personnel. https://doi.org/10.1371/journal.pone.0196562
  17. Milkias Barkot, Ametaj Amantia, Alemayehu Melkam et al. (2022). Psychometric properties and factor structure of the Kessler-10 among Ethiopian adults. https://doi.org/10.1016/j.jad.2022.02.013
  18. Naisanga Molly, Ametaj Amantia, Kim Hannah H et al. (2022). Construct validity and factor structure of the K-10 among Ugandan adults. https://doi.org/10.1016/j.jad.2022.05.022
  19. Hoffman Jacob, Cossie Qhama, Ametaj Amantia A et al. (2022). Construct validity and factor structure of the Kessler-10 in South Africa. https://doi.org/10.1186/s40359-022-00883-9
  20. Osman Naweed, Chow Winnie S, Michel Chantal et al. (2022). Psychometric properties of the Kessler psychological scales in a Swiss young-adult community sample indicate poor suitability for community screening for mental disorders. https://doi.org/10.1111/eip.13296
  21. Lantos Dorottya, Moreno-Agostino Darío, Harris Lasana T et al. (2023). The performance of long vs. short questionnaire-based measures of depression, anxiety, and psychological distress among UK adults: A comparison of the patient health questionnaires, generalized anxiety disorder scales, malaise inventory, and Kessler scales. https://doi.org/10.1016/j.jad.2023.06.033
  22. Chamberlain Peter, Goldney Robert, Delfabbro Paul et al. (2009). Suicidal ideation. The clinical utility of the K10. https://doi.org/10.1027/0227-5910.30.1.39
  23. Burns Kristy, Schroeder Elizabeth-Ann, Fung Thomas et al. (2023). Industry differences in psychological distress and distress-related productivity loss: A cross-sectional study of Australian workers. https://doi.org/10.1002/1348-9585.12428
  24. Laditka James N, Laditka Sarah B, Arif Ahmed A et al. (2023). Psychological distress is more common in some occupations and increases with job tenure: a thirty-seven year panel study in the United States. https://doi.org/10.1186/s40359-023-01119-0
  25. Ul Husnain Muhammad Iftikhar, Hajizadeh Mohammad, Ahmad Hasnat et al. (2024). The Hidden Toll of Psychological Distress in Australian Adults and Its Impact on Health-Related Quality of Life Measured as Health State Utilities. https://doi.org/10.1007/s40258-024-00879-z
  26. Pereira Anabela, Oliveira Carla Andreia, Bártolo Ana et al. (2019). Reliability and Factor Structure of the 10-item Kessler Psychological Distress Scale (K10) among Portuguese adults. https://doi.org/10.1590/1413-81232018243.06322017
  27. Easton Scott D, Safadi Nadia S, Wang Yao (2017). The Kessler psychological distress scale: translation and validation of an Arabic version https://doi.org/10.1186/s12955-017-0783-9
  28. Peixoto Evandro Morais, Zanini Daniela Sacramento, de Andrade Josemberg Moura (2021). Cross-cultural adaptation and psychometric properties of the Kessler Distress Scale (K10): an application of the rating scale model https://doi.org/10.1186/s41155-021-00186-9
  29. Inoue Yosuke, Hikichi Hiroyuki, Inoue Mariko et al. (2022). Workplace Social Support and Reduced Psychological Distress: A 1-Year Occupational Cohort Study https://doi.org/10.1097/JOM.0000000000002675
  30. Perrelli Jaqueline Galdino Albuquerque, Vasconcelos Gabriel Vinicius Souza de, Correia e Sa Jessica Rodrigues et al. (2024). Validity of the Kessler Psychological Distress scale in Brazilian higher education students https://doi.org/10.1590/1518-8345.7073.4254

Record notes

Overall confidence: the K10 is a strong screener for the likely presence of common mental disorders against a diagnostic reference standard (High, well-established, principally Australian, US and Canadian samples), with high internal consistency (High, well-established) and extensive norms (High, but not UK-specific for working adults). Its factor structure is genuinely contested (recorded as contested, not averaged), and three properties are honestly Absent or thin: test-retest reliability (no coefficient located in the latest review pass), responsiveness/MIC (none located), and organisational criterion validity (no work-outcome validation exists; the workplace literature uses the K10/K6 as an exposure or prevalence measure, not as a validated work-outcome screen). For the UK working-adult audience every graded property is indirect: the evidence was earned in non-UK, non-workplace, general-population, student, clinical or military samples, and the instrument is clinical/epidemiological in origin, making any workplace deployment off-label (see deployment_context_caveat). Licence verified against the Harvard Medical School NCS steward page in the latest review pass (2026-07-12): free to use, no permission or fee required, cite the source; no named open-content licence code is asserted by the steward. A few supporting sources are named in prose but not separately DOI-listed where they duplicate an already-cited finding (Arabic validation Easton 2017 doi:10.1186/s12955-017-0783-9; Brazilian rating-scale-model adaptation, Peixoto 2021 doi:10.1186/s41155-021-00186-9; Japanese workplace social-support cohort using K6 doi:10.1097/JOM.0000000000002675); these resolve but were retrieved as context rather than as primary graded evidence. Schema v0.2 friction: the criterion split worked well and made the organisational Absent finding explicit; the single evidence_form tag per property is a simplification where evidence is genuinely mixed (canonical K10 plus K6 short-form and parent-item evidence), flagged as 'mixed' where that applies.