Patient Health Questionnaire-9 (PHQ-9)
Identity
Version: PHQ-9 (2001); abbreviated variants PHQ-8 (omits item 9) and PHQ-2 (first two items) are in wide use
Structure: 9 items, each scored 0 to 3 (total 0 to 27); severity bands 5, 10, 15, 20 map to mild, moderate, moderately severe and severe
Original citation: Kroenke K, Spitzer RL, Williams JB (2001) The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine. 10.1046/j.1525-1497.2001.016009606.x
Steward / publisher: Developed by Spitzer, Williams and Kroenke as the depression module of the PRIME-MD PHQ, under an educational grant from Pfizer Inc; the PHQ suite was subsequently released by Pfizer for free public access.
Constructs claimed
Severity of depressive symptoms over the preceding two weeks and, secondarily, provisional detection of major depressive disorder. The nine items map directly onto the nine DSM-IV/DSM-5 symptom criteria for a major depressive episode, so the instrument claims both a continuous severity construct and a criterion-referenced screening function (Kroenke 2001). It is a depression-specific measure, not a general wellbeing or distress index, and was designed for primary-care and clinical use rather than for occupational or workplace-wellbeing measurement.
Evidence
Structural validity Highcontestedevidence form: canonical
indirectness see findings
The dimensionality of the PHQ-9 is genuinely contested, but the practical reading across studies is that it behaves as an essentially unidimensional depression severity scale even where a two-factor model fits marginally better. The original development treated the total score as a single severity dimension (Kroenke 2001). Within the individual participant data (IPD) programme, a comparison of unidimensional, two-dimensional (cognitive/affective versus somatic) and bifactor latent models found that all fitted reasonably and that scoring by the more complex latent models improved sensitivity only marginally (by about 0.04 to 0.05) while reducing specificity, relative to the sum score at a cut-off of 10 or above (Fischer et al 2021). Several population studies reach the same conclusion from different angles: in Korean nationally representative data a single-factor model fitted well (CFI 0.944) with satisfactory internal consistency (Lee et al 2023); in Brazilian university students both one- and two-dimensional models fitted, with the authors favouring the unidimensional solution (Rufino et al 2024); and after traumatic brain injury a general factor explained about 85% of the variance despite a slight bifactor improvement (Teymoori et al 2020). Where a two-factor (somatic versus cognitive-affective) structure is reported, the factors tend to be very highly correlated: in a stroke sample the two-factor model fitted slightly better than the one-factor model (CFI 0.984 versus 0.974) but the inter-factor correlation of 0.866 pointed back to unidimensionality (Blake et al 2024), and an Italian coronary heart disease sample reported a bi-dimensional somatic/cognitive solution (Di et al 2025). The disagreement is therefore about whether the somatic items form a distinguishable subfactor, not about whether a single severity score is defensible; the weight of evidence supports scoring and interpreting a single total.
Confidence note: High: many good-quality factor-analytic studies across large and varied samples, converging on essentially unidimensional use despite a recurring, well-characterised somatic/cognitive two-factor debate.
Convergent and discriminant validity Moderatecontestedevidence form: canonical
indirectness see findings
Convergent evidence is consistent in direction but uneven in magnitude, and discriminant separation from anxiety is imperfect. In development, higher PHQ-9 scores tracked substantial decrements across all six SF-20 functional-status subscales and greater symptom-related difficulty, supporting construct validity (Kroenke 2001). In the Chinese general population the PHQ-9 correlated negatively with SF-36 subscales (r from -0.11 to -0.47) as expected, but its correlation with the Zung Self-Rating Depression Scale was unexpectedly weak (r = 0.29), an inconsistency worth noting against the usual assumption of strong convergence with other depression measures (Wang et al 2014). Convergent relationships with sleep quality, alcohol use and physical activity have also been reported in students (Rufino et al 2024). On discriminant validity, the PHQ-9 and the GAD-7 anxiety scale share a large common factor: after traumatic brain injury a general distress factor dominated both scales (about 85% of variance) even though the instruments related differently to SF-36 subscales (Teymoori et al 2020), and in coronary heart disease PHQ-9 and GAD-7 scores were significantly positively correlated (Di et al 2025). Depression and anxiety symptoms measured this way are statistically distinguishable but strongly overlapping.
Confidence note: Moderate: multiple studies establish expected convergent and functional-status correlations, but magnitudes vary (including one weak convergent correlation) and discriminant separation from anxiety is only partial.
Criterion validity: reference standard Highwell-establishedevidence form: canonical
indirectness see findings
Criterion validity against a diagnostic interview for major depression is the PHQ-9's best-evidenced property, and it is strong; criterion validity against organisational or workplace outcomes is, by contrast, not established in the retrieved literature. In the original clinical validation a cut-off of 10 or above gave 88% sensitivity and 88% specificity against an independent mental-health-professional interview (Kroenke 2001). The individual participant data meta-analysis programme is the anchor here. The 2019 IPD meta-analysis (58 studies, n = 17,357, 2,312 major depression cases) found combined sensitivity and specificity maximised at a cut-off of 10 or above against semistructured interviews (sensitivity 0.88, 95% CI 0.83 to 0.92; specificity 0.85, 0.82 to 0.88) (Levis et al 2019). The 2021 update (100 studies, n = 44,503) reproduced this at the same cut-off (sensitivity 0.85, 0.79 to 0.89; specificity 0.85, 0.82 to 0.87) (Negeri et al 2021). A crucial nuance is reference-standard dependence: sensitivity against semistructured clinician interviews ran markedly higher than against fully structured lay interviews (median difference around 21%) or the MINI (around 11%), while specificity was similar across standards (Levis et al 2019, Negeri et al 2021). The diagnostic algorithm approach performed worse than the cut-off (algorithm sensitivity around 0.57 to 0.61 versus 0.88 for the cut-off against semistructured interviews) (He et al 2019), and the PHQ-2 with cut-off 3 or above (sensitivity 0.72, specificity 0.85) is a reasonable first-stage screen followed by the PHQ-9 (Levis et al 2020). For the health-related associations closest to organisational outcomes, the development study showed that self-reported sick days and health-care utilisation rose with PHQ-9 severity (Kroenke 2001), but no study retrieved in the latest review pass tested PHQ-9 scores against measured sickness absence, staff turnover, productivity loss or other workplace outcomes as criteria. Criterion validity against organisational endpoints should therefore be treated as unestablished.
Confidence note: High for diagnostic criterion validity against clinical interviews (large, consistent IPD evidence); Absent for criterion validity against organisational/workplace outcomes (no such studies retrieved). The overall grade is split by criterion.
Criterion validity: organisational Absent (a finding about the literature)untestedevidence form: canonical
indirectness see findings
Organisational criterion evidence (sickness absence, turnover, performance, diagnosed conditions in a work context): see the criterion findings; graded from the pass-one record.
Confidence note: High for diagnostic criterion validity against clinical interviews (large, consistent IPD evidence); Absent for criterion validity against organisational/workplace outcomes (no such studies retrieved). The overall grade is split by criterion.
Internal consistency Highwell-establishedevidence form: canonical
indirectness see findings
Internal consistency is consistently high across populations. A reliability generalisation meta-analysis of 60 studies (232,147 participants) estimated a pooled Cronbach's alpha of 0.86 (95% CI 0.85 to 0.87), with self-administered formats slightly higher (alpha 0.87) than face-to-face interview administration (alpha 0.80), though between-study heterogeneity was very large (I-squared 99.3%) (Ajele 2025). Single-study estimates sit in the same range: alpha 0.86 in the Chinese general population (Wang et al 2014) and omega 0.812 in a Korean nationally representative sample (Lee et al 2023). A systematic review of the wider PHQ family likewise reported good internal consistency for the PHQ-9 (Kroenke 2010), and a meta-analysis of the Spanish-language versions evaluated internal consistency alongside accuracy (Martinez et al 2023). Reported coefficients of omega tend to align with alpha where both are given, but omega is reported far less often than alpha.
Confidence note: High: a large reliability-generalisation meta-analysis plus multiple primary studies converge on alpha around 0.86, albeit with substantial heterogeneity and sparse omega reporting.
Test-retest reliability Moderatethin
indirectness see summary
Test-retest reliability exists but is markedly under-studied relative to the vast diagnostic-accuracy literature, and this asymmetry is itself the finding. The reliability generalisation meta-analysis could pool a test-retest estimate from only 8 of its 60 studies, yielding 0.82 (95% CI 0.74 to 0.90) (Ajele 2025). Primary estimates are favourable where they exist: a two-week retest correlation of 0.86 in the Chinese general population (Wang et al 2014), test-retest described as excellent over a seven-day interval in a late-life depression treatment sample (Lowe et al 2004), and intraclass agreement assessed for self- versus telephone-administration in Spanish primary care (Pinto-Meza et al 2005). No test-retest study in an occupational or workplace sample, and none using a UK working population, was located in this pass. The property is present and reassuring in the settings studied, but the evidence base is thin and skewed toward clinical and general-population samples over short intervals.
Confidence note: Moderate: several favourable estimates (roughly 0.82 to 0.86) exist and a meta-analytic pooled value is available, but from few studies (n = 8 pooled), short intervals, and no workplace or UK occupational data. Not absent, but comparatively neglected.
Measurement invariance Lowthinevidence form: canonical
indirectness see findings
Invariance has been tested piecemeal and mostly reaches metric to scalar level within the groups examined, but coverage of occupational and UK working populations is absent. In a Korean nationally representative sample the one-factor PHQ-9 showed equivalent structure, factor loadings and item intercepts across age groups, i.e. up to scalar invariance across ages (Lee et al 2023). Invariance across sex and across age (65 and over versus under 65) was tested in an Italian coronary heart disease cohort (Di et al 2025). Item response theory analysis in a large Danish implantable-defibrillator cohort found no differential item functioning across educational level, age, clinical indication or heart-failure severity, with only a single item showing DIF by gender (Pedersen et al 2016). This points to broadly stable measurement across sex and age in clinical and general populations. However, the studies are confined to clinical or general-population samples; invariance across occupations, across employed versus unemployed status, over repeated workplace administrations, and within UK working populations was not established in retrieved evidence. Longitudinal (over-time) invariance is likewise weakly evidenced for the PHQ-9 specifically.
Confidence note: Low to Moderate: scalar invariance is demonstrated across age (one strong study) and DIF is minimal in another, but evidence is scattered across clinical populations, occupation and UK-workplace invariance is untested, and over-time invariance is weak.
Responsiveness and MIC Moderatethinevidence form: canonical
indirectness see findings
Responsiveness to change is well supported, and a minimal important change has been proposed, though the MIC rests largely on one study. In the IMPACT late-life depression trial (n = 434 intervention participants) the PHQ-9 was responsive to treatment, with an effect size (about -1.3 at three months) exceeding the SCL-20 depression scale at three months and matching it at six months; change scores discriminated persistent depression, partial remission and full remission against structured diagnostic interviews (Lowe et al 2004). That study estimated a minimal clinically important difference for individual change of about 5 points on the 0 to 27 scale, derived as two standard errors of measurement (Lowe et al 2004). A systematic review of the PHQ family concluded that sensitivity to change is well established for the PHQ-9 (Kroenke 2010). Routine-outcome use in stepped-care services also relies implicitly on responsiveness, with large pre-post effect sizes reported for depression (Richards 2009). The MIC of 5 points is widely cited but should be understood as a single-derivation estimate rather than a triangulated consensus value.
Confidence note: Moderate: responsiveness is demonstrated in a good treatment study and endorsed by a review, but the 5-point MIC derives essentially from one study and one method (2 SEM), and no MIC has been established in a workplace context.
Populations, languages and norms
The PHQ-9 has been validated across a wide range of clinical, general-population and disease-specific samples and in many languages, but formal population norms in the classical sense are scarce; the UK benchmark is embedded in a national service dataset rather than a norm table. The IPD meta-analyses aggregate around 100 primary studies from many countries (Negeri et al 2021, Levis et al 2020). Validated non-English versions retrieved in the latest review pass include Chinese (Wang et al 2014) and Spanish, the latter via a dedicated systematic review and meta-analysis of the Spanish-language PHQ-2 and PHQ-9 (Martinez et al 2023), with further evaluations in Korean (Lee et al 2023), Danish (Pedersen et al 2016) and Italian (Di et al 2025) samples. For the UK specifically, the PHQ-9 is the routine depression outcome measure of the English NHS Talking Therapies programme (formerly Improving Access to Psychological Therapies, IAPT), where a score of 10 or above defines 'caseness' and movement below it underpins recovery reporting; large UK service cohorts and a UK randomised trial have used it in exactly this way (Richards 2009, Barkham et al 2021). UK reference values therefore live in national service reporting rather than in a published normative sample. General-population norms exist chiefly through large national surveys used in the invariance literature rather than as a dedicated UK norm set.
Criticisms and controversies
Several substantive controversies recur in the literature. First, cut-point and accuracy estimation. Studies that selectively report only well-performing cut-offs bias meta-analytic accuracy: for the PHQ-9, published results underestimated sensitivity below a cut-off of 10 (median difference about -0.06) and overestimated it above 10 (median about +0.07) (Neupane et al 2021). Relatedly, using small datasets to simultaneously choose an optimal cut-off and estimate its accuracy is biased; in resampling from the IPD database the population-level optimal cut-off was actually 8 or above, and small studies scattered widely around it (only about 17% of 100-participant studies recovered the true optimum) (Levis et al 2024). This questions the near-universal reliance on a cut-off of exactly 10. Second, reference-standard dependence: reported sensitivity is substantially higher against semistructured clinician interviews than against fully structured lay interviews or the MINI, so headline accuracy figures are conditional on the comparator (Levis et al 2019, Negeri et al 2021, He et al 2019). Third, somatic-item confounding in physically ill populations. In systemic sclerosis, somatic items accounted for a larger share of the total score than in matched healthy respondents, inflating scores by roughly 1.0 to 1.4 points (Hedges g 0.38 to 0.55) (Leavens et al 2012); in stroke, summed scoring moderately overestimated depression relative to a comparison sample (Cohen d about 0.43), driven by items such as tiredness and appetite (Blake et al 2024). This matters wherever respondents have physical illness or fatigue. Fourth, item 9 (thoughts of death or self-harm) is often misread as a suicidality measure, yet most positive responses are not associated with suicidality; the PHQ-8, which omits item 9, correlates almost perfectly with the PHQ-9 (r = 0.996) and performs almost identically for detecting depression (Wu et al 2019). Fifth, complexity does not pay: latent-variable and machine-learning scoring add negligible accuracy over the simple sum score with a cut-off (Fischer et al 2021, Hong et al 2022). Finally, and central to the OWHS use case, the PHQ-9 is a clinical depression screener whose criterion validity was earned in primary-care and clinical populations against diagnostic interviews. Workplace deployment is a different context: screening in a lower-prevalence, largely non-help-seeking workforce reduces positive predictive value, the somatic-confounding problem is relevant to occupational groups with physical demands or illness, and no retrieved study validates the PHQ-9 against workplace criteria. The instrument's clinical provenance must not be read as endorsement of clinical-grade performance in a workplace-wellbeing programme; deployments that used it in worker samples treated it as an off-the-shelf symptom measure rather than validating it there (Doki et al 2024).
References (26)
- Kroenke K, Spitzer RL, Williams JB (2001). The PHQ-9: validity of a brief depression severity measure https://doi.org/10.1046/j.1525-1497.2001.016009606.x
- Kroenke K, Spitzer RL, Williams JB, Lowe B (2010). The Patient Health Questionnaire Somatic, Anxiety, and Depressive Symptom Scales: a systematic review https://doi.org/10.1016/j.genhosppsych.2010.03.006
- Levis B, Benedetti A, Thombs BD, et al (2019). Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis https://doi.org/10.1136/bmj.l1476
- Negeri ZF, Levis B, Sun Y, et al (2021). Accuracy of the Patient Health Questionnaire-9 for screening to detect major depression: updated systematic review and individual participant data meta-analysis https://doi.org/10.1136/bmj.n2183
- Wu Y, Levis B, Riehm KE, et al (2019). Equivalency of the diagnostic accuracy of the PHQ-8 and PHQ-9: a systematic review and individual participant data meta-analysis https://doi.org/10.1017/S0033291719001314
- He C, Levis B, Riehm KE, et al (2019). The Accuracy of the Patient Health Questionnaire-9 Algorithm for Screening to Detect Major Depression: An Individual Participant Data Meta-Analysis https://doi.org/10.1159/000502294
- Levis B, Sun Y, He C, et al (2020). Accuracy of the PHQ-2 Alone and in Combination With the PHQ-9 for Screening to Detect Major Depression https://doi.org/10.1001/jama.2020.6504
- Neupane D, Levis B, Bhandari PM, et al (2021). Selective cutoff reporting in studies of the accuracy of the Patient Health Questionnaire-9 and Edinburgh Postnatal Depression Scale https://doi.org/10.1002/mpr.1873
- Levis B, Bhandari PM, Neupane D, et al (2024). Data-Driven Cutoff Selection for the Patient Health Questionnaire-9 Depression Screening Tool https://doi.org/10.1001/jamanetworkopen.2024.29630
- Fischer F, Levis B, Falk C, et al (2021). Comparison of different scoring methods based on latent variable models of the PHQ-9: an individual participant data meta-analysis https://doi.org/10.1017/S0033291721000131
- Hong ZM, Williams J, Bulloch A, et al (2022). Alternative scoring of the Patient Health Questionnaire-9 in neurological populations: an approach based on a predictive algorithm deriving from individual item scores https://doi.org/10.1016/j.genhosppsych.2022.04.011
- Rufino JV, Rodrigues R, Birolim MM, et al (2024). Analysis of the dimensional structure of the Patient Health Questionnaire-9 (PHQ-9) in undergraduate students at a public university in Brazil https://doi.org/10.1016/j.jad.2024.01.051
- Blake JJ, Munyombwe T, Fischer F, et al (2024). The factor structure of the Patient Health Questionnaire-9 in stroke: A comparison with a non-stroke population https://doi.org/10.1016/j.jpsychores.2024.111983
- Teymoori A, Gorbunova A, Haghish FE, et al (2020). Factorial Structure and Validity of Depression (PHQ-9) and Anxiety (GAD-7) Scales after Traumatic Brain Injury https://doi.org/10.3390/jcm9030873
- Wang W, Bian Q, Zhao Y, et al (2014). Reliability and validity of the Chinese version of the Patient Health Questionnaire (PHQ-9) in the general population https://doi.org/10.1016/j.genhosppsych.2014.05.021
- Lowe B, Unutzer J, Callahan CM, et al (2004). Monitoring depression treatment outcomes with the Patient Health Questionnaire-9 https://doi.org/10.1097/00005650-200412000-00006
- Leavens A, Patten SB, Hudson M, et al (2012). Influence of somatic symptoms on Patient Health Questionnaire-9 depression scores among patients with systemic sclerosis compared to a healthy general population sample https://doi.org/10.1002/acr.21675
- Lee EH, Kang EH, Kang HJ, et al (2023). Measurement invariance of the patient health questionnaire-9 depression scale in a nationally representative population-based sample https://doi.org/10.3389/fpsyg.2023.1217038
- Di Matteo R, Bolgeo T, Simonelli N, et al (2025). Psychometric Properties and Measurement Invariance of the Patient Health Questionnaire 9 in an Italian Coronary Heart Disease Population https://doi.org/10.1097/JCN.0000000000001178
- Pedersen SS, Mathiasen K, Christensen KB, et al (2016). Psychometric analysis of the Patient Health Questionnaire in Danish patients with an implantable cardioverter defibrillator (The DEFIB-WOMEN study) https://doi.org/10.1016/j.jpsychores.2016.09.010
- Ajele KW, Idemudia ES (2025). Charting the course of depression care: a meta-analysis of reliability generalization of the Patient Health Questionnaire (PHQ-9) as the measure https://doi.org/10.1007/s44192-025-00181-x
- Martinez A, Teklu SM, Tahir P, et al (2023). Validity of the Spanish-Language Patient Health Questionnaires 2 and 9: A Systematic Review and Meta-Analysis https://doi.org/10.1001/jamanetworkopen.2023.36529
- Pinto-Meza A, Serrano-Blanco A, Penarrubia MT, et al (2005). Assessing depression in primary care with the PHQ-9: can it be carried out over the telephone? https://doi.org/10.1111/j.1525-1497.2005.0144.x
- Barkham M, Saxon D, Hardy GE, et al (2021). Person-centred experiential therapy versus cognitive behavioural therapy delivered in the English Improving Access to Psychological Therapies service for the treatment of moderate or severe depression (PRaCTICED) https://doi.org/10.1016/S2215-0366(21)00083-3
- Richards DA, Suckling R (2009). Improving access to psychological therapies: phase IV prospective cohort study https://doi.org/10.1348/014466509X405178
- Doki S, Hori D, Takahashi T, et al (2024). Designing a test battery for workers' well-being: the first wave of the Tsukuba Salutogenic Occupational Cohort Study https://doi.org/10.1265/ehpm.23-00372
Record notes
[Upgraded from v0.1 to v0.2 structure in pass two; criterion field split, licence re-verified 2026-07-12.] Overall confidence: the PHQ-9 is the confidence-grading high-water mark. Structural validity, diagnostic criterion validity and internal consistency are High, anchored on the Levis/Thombs individual participant data programme and a reliability-generalisation meta-analysis; responsiveness and the somatic-confounding and cut-point critiques are well evidenced. The genuinely weaker or absent cells are test-retest (present but from few studies, none occupational), measurement invariance (scattered, no occupational or UK-workplace coverage, weak over-time evidence), criterion validity against organisational outcomes (absent), and workplace-specific psychometrics (absent). Schema stress-test notes: (1) The single 'criterion_validity' field conflates two very different evidence bases, diagnostic criterion validity (High) versus criterion validity against organisational outcomes (Absent). The maintainers graded it as split and said so in the findings, but a schema that forced one grade would misrepresent the instrument; the field should ideally be divided. (2) 'confidence' is a single scalar per property, yet for several properties the honest grade differs by sub-question (e.g. invariance is scalar-level across age but untested across occupation). The maintainers encoded the dominant grade and qualified it in the justification. (3) The clinical-origin versus workplace-deployment gap is the most important caveat for this registry and does not have a dedicated field; The maintainers carried it in criticisms_controversies and flagged it in constructs_claimed and criterion_validity, but it risks being lost if a reader scans only the property grades. (4) Licence status is factually clear (public domain, Pfizer-released) but the maintainers could not attach a DOI-bearing primary source to the licensing act itself, only to the original validation paper; The maintainers flagged this rather than attach a non-resolvable citation. (5) 'Absent' was used strictly for organisational criterion validity and workplace psychometrics; test-retest was deliberately NOT graded Absent because evidence exists, only sparsely, which the scale's wording ('barely-studied') made a close call between Low and Moderate.