multi-item-scale · reviewed 2026-07-12 contested: Structural validity

GAD-7 (Generalised Anxiety Disorder-7)

Licence verified: 2026-07-12 · record reviewed: 2026-07-12 (pass two)

Identity

Version: Original 7-item GAD-7 (2006); a 2-item short form (GAD-2, first two items) is also distributed

Structure: 7

Original citation: Spitzer RL, Kroenke K, Williams JBW, Lowe B. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med. 2006;166(10):1092-1097. doi:10.1001/archinte.166.10.1092

Steward / publisher: Developed by Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues with an educational grant from Pfizer Inc.; distributed by the steward site phqscreeners.com (copyright Pfizer Inc.)

Licence status (verified 2026-07-12): Free to use, no permission required. The steward's currently distributed GAD-7 English PDF (phqscreeners.com) carries the footer 'No permission required to reproduce, translate, display or distribute', developed with an educational grant from Pfizer Inc.; the steward instrument entry also carries 'Copyright (c) Pfizer Inc. All rights reserved.' This is a no-permission-required grant with copyright retained by Pfizer, NOT a formal open or Creative Commons licence. Confirmed at the last verification attempt from the steward's distributed PDF footer text (body content, not merely page titles) and the LOINC mirror of the steward entry; direct HTML fetches of phqscreeners.com returned HTTP 403.
Source: Verbatim footer text of the official GAD-7 English PDF hosted on the steward site (phqscreeners.com/images/sites/g/files/g10060481/f/201412/GAD-7_English.pdf), read in the latest review pass: 'Developed by Drs. Robert L. Spitzer, Janet B.W. Williams, Kurt Kroenke and colleagues, with an educational grant from Pfizer Inc. No permission required to reproduce, translate, display or distribute.' The steward instrument entry (as mirrored by LOINC 69737-5, which cites URL https://www.phqscreeners.com/) additionally records 'Copyright (c) Pfizer Inc. All rights reserved.' alongside the same no-permission-required statement. Direct urllib fetches of phqscreeners.com pages returned HTTP 403 (server-side bot refusal, not a sandbox block), so the current terms were confirmed from the steward's own distributed PDF footer text and the LOINC mirror of the steward entry rather than the rendered HTML page. Not sourced from the founding paper.

Constructs claimed

Severity of generalised anxiety disorder symptoms over the preceding two weeks. Functions both as a screener for probable GAD and as a continuous self-report measure of anxiety-symptom severity.

Evidence

Deployment context caveat. GAD-7 is a clinical screener developed and validated in primary care against a psychiatric diagnostic interview (Spitzer 2006, doi:10.1001/archinte.166.10.1092). Any workplace or organisational deployment is a different context from its validation setting. It measures anxiety-symptom severity and screens for probable GAD; it is not a diagnostic instrument and not a fitness-for-work measure. A positive screen indicates a need for clinical assessment, never an employment or performance decision. Its clinical origin may be stated as fact, but its psychometric performance in a workplace deployment context should not be assumed from clinical/primary-care evidence. (applies to every property below)

Structural validity Highcontestedevidence form: canonical

direct unidimensional evidence includes UK-inclusive nationally representative samples (Shevlin 2022), though much replication is in non-UK, clinical and disease-specific populations

GAD-7 was designed as, and is most often confirmed as, a single-factor (unidimensional) scale. The original development study confirmed anxiety and depression as distinct dimensions (Spitzer 2006), and a nationally representative German general-population confirmatory factor analysis (N=5030) substantiated a one-dimensional structure with factorial invariance for gender and age (Lowe 2008). Unidimensional fit has been replicated across many settings: Cypriot perinatal women (Vogazianos 2022), an Italian coronary heart disease sample (Bolgeo 2023), four nationally representative European samples including the UK (Shevlin 2022), Canadian young adults where a one-factor model fit best (Riglea 2025), and 20 Czech samples (N=5529) supporting a unidimensional structure (Cigler 2026). However, a competing two-factor structure separating cognitive-emotional from somatic items recurs: a network analysis of Spanish primary-care patients revealed a two-factor solution (Moriana 2021), a US college-student study fit both one- and two-factor models (White 2025), a large psychotherapy sample found the scale technically multidimensional though sum scores remained justifiable (Stochl 2020), and a Malaysian study preferred a six-item second-order model (Pheh 2023). The one-factor model dominates practice and fits well in most samples, but the dimensionality question is genuinely unsettled.

Sub-grades (evidence differs by subgroup):

  • {"subgroup": "one-factor (unidimensional) model", "grade": "High", "note": "Well replicated across languages and settings; the dominant and recommended scoring model"}
  • {"subgroup": "cognitive-emotional vs somatic two-factor model", "grade": "Moderate", "note": "Recurs in network analyses and some CFAs; keeps dimensionality contested rather than settled"}

Convergent and discriminant validity Moderatewell-establishedevidence form: canonical

indirect convergent/discriminant evidence is mostly from general-population, student and clinical non-workplace samples (German general population, US students, disease-specific cohorts)

Convergent validity is consistently supported. In the German general population GAD-7 correlated r=0.64 with the PHQ-2 depression module and r=-0.43 with the Rosenberg Self-Esteem Scale (Lowe 2008). In US college students the total score correlated r=0.70 with the trait scale of the State-Trait Anxiety Inventory (convergent) and only r=-0.04 with a behavioural activation reward subscale (discriminant) (White 2025). Increasing scores were strongly associated with multiple domains of functional impairment in the original study (Spitzer 2006). The recurring discriminant concern is the strong overlap with depression: GAD-7 and PHQ-9 anxiety and depression factors are highly correlated and frequently co-occur (Stochl 2020; Bolgeo 2023), so the scale distinguishes anxiety from unrelated constructs well but discriminates anxiety from depression less cleanly.

Criterion validity: reference standard Highwell-establishedevidence form: canonical

indirect validated against diagnostic interview in primary-care and disease-specific clinical samples, not in workplace populations, and UK-specific diagnostic-accuracy data are a minority of the pooled evidence

Against a structured or semi-structured clinical interview, GAD-7 has been extensively evaluated. The original primary-care study identified a cut-off of 10 with sensitivity 89% and specificity 82% for GAD (Spitzer 2006). An early diagnostic meta-analysis (12 samples, 5223 participants) found pooled sensitivity 0.83 and specificity 0.84 at a cut-off of 8, with cut-offs 7 to 10 performing similarly (Plummer 2016). The most comprehensive synthesis, a 2025 Cochrane review of 48 studies (19,228 participants, 27 countries, 24 languages), reported that at the recommended cut-off of 10 or higher the GAD-7 summary sensitivity was 0.64 (95% CI 0.56 to 0.72) and specificity 0.91 (95% CI 0.87 to 0.93) for detecting GAD, with an area under the curve of 0.86; for detecting any anxiety disorder sensitivity fell to 0.48 (specificity 0.91) (Akturk 2025). The pooled sensitivity at cut-off 10 is therefore markedly lower than the original single-study estimate, with pronounced heterogeneity, and the Cochrane authors caution that the summary estimates are rough averages that may deviate substantially in specific situations. Setting-specific validations against interview (for example Taiwanese epilepsy patients, optimal cut-off 7) add further threshold variability (Shih 2022).

Criterion validity: organisational Absent (a finding about the literature)untestedevidence form: mixed

indirect the closest evidence uses the derivative GAD-2 and self-reported work outcomes, and no study links the full GAD-7 to objective organisational records

No study validating the full GAD-7 against objective organisational outcomes (recorded sickness absence, turnover, or measured job performance) was located in the latest review pass. The nearest evidence is adjacent rather than direct: a study of 4953 working Australians linked probable anxiety to worse self-reported presenteeism and absenteeism on the WHO Health and Work Performance Questionnaire, but it used the 2-item GAD-2, not the full GAD-7, and relied on self-reported rather than employer-recorded outcomes (Deady 2021). A Polish validation among employees related GAD-7 to professional burnout and psychological distress, again self-reported constructs rather than organisational records (Basinska 2023). The original study related GAD-7 to self-reported disability days and functional impairment, not to work-context criterion outcomes (Spitzer 2006). Organisational criterion validity for the fielded GAD-7 is therefore essentially untested.

Internal consistency Highwell-establishedevidence form: canonical

direct UK-inclusive nationally representative samples contribute (Shevlin 2022, Saunders 2023), although most individual alpha estimates come from non-UK or clinical samples

Internal consistency is uniformly high across populations and languages. Cronbach's alpha was 0.89, identical across all gender and age subgroups, in the German general population (Lowe 2008); 0.91 in US college students (White 2025); alpha 0.907 with McDonald's omega 0.909 in Cypriot perinatal women (Vogazianos 2022); alpha 0.89 with composite reliability 0.90 in an Italian cardiac sample (Bolgeo 2023); and 0.928 in Taiwanese epilepsy patients (Shih 2022). Lower but still acceptable values appear in some translations, for example alpha 0.81 in a Swahili HIV sample (Nyongesa 2020) and a median alpha of 0.86 across 20 Czech samples (Cigler 2026). Values consistently sit in the 0.81 to 0.93 range.

Test-retest reliability Lowthin

indirect the original ICC of 0.83 is reported second-hand in retrieved sources without interval or sample size, and the two dedicated coefficients located are from a Kenyan HIV sample and a Czech life-events design, none in a UK working population

CoefficientTypeIntervalSamplePopulationEvidence form
0.83ICCnot reported in retrieved sourcesnot reported in retrieved sourcesUS primary care (original validation)canonical
0.59ICC2 weeks60Adults living with HIV, Kilifi, Kenya (Swahili version)canonical
0.46 to 0.53rrepeated measures over 2 to 4 time points spanning major life eventssubset of 5529Czech general-population and psychiatric adultscanonical

Dedicated short-interval test-retest studies of GAD-7 are sparse. The original validation reportedly gave ICC 0.83, but its interval and sample size were not captured in sources retrieved in the latest review pass (the value is cited second-hand in Nyongesa 2020, doi:10.1186/s12991-020-00312-4). A 2-week ICC of 0.59 was found in a Swahili HIV sample (Nyongesa 2020), and the Czech multi-sample study found moderate stability of r=0.46 to 0.53, but over intervals spanning major life events rather than a fixed short retest window (Cigler 2026, doi:10.1016/j.janxdis.2026.103149). Test-retest reliability for a UK working-adult population is untested. The absence of a consistent, short-interval, well-sampled retest estimate is itself the finding.

Measurement invariance Highwell-establishedevidence form: canonical

direct UK treatment-seeking and working-age-versus-older samples are included (Saunders 2023, Delamain 2024), though occupational-group invariance specifically is untested

Measurement invariance is one of the most heavily studied properties of the GAD-7, and it is largely supported across groups but contested over time. Across sex/gender, invariance or an absence of differential item functioning is repeatedly demonstrated, including in very large UK treatment-seeking samples (N=165,872) (Saunders 2023), Canadian adolescents where strict invariance held by sex and grade (Romano 2021), and four European nationally representative samples with no DIF across sex, age or country (Shevlin 2022). Across age, invariance held between UK working-age and older adults with only limited DIF (Delamain 2024). Across language and country, invariance is supported across English and French (Riglea 2025), across 18 countries after traumatic brain injury (Teymoori 2020), and with full scalar invariance in rural India (De Man 2021). Longitudinal invariance is where the literature conflicts: strict temporal invariance was established across 10 psychotherapy sessions (Stochl 2020) and across time in Czech and Canadian samples (Cigler 2026; Riglea 2025), yet longitudinal invariance was NOT established in a partial-hospital sample, implying that raw pre-post change scores may be unreliable in that setting (Ong 2021).

Sub-grades (evidence differs by subgroup):

  • {"subgroup": "sex / gender", "grade": "High", "note": "Strongly supported across many samples including large UK datasets"}
  • {"subgroup": "age", "grade": "High", "note": "Supported including UK working-age versus older adults (Delamain 2024)"}
  • {"subgroup": "language / country", "grade": "Moderate", "note": "Supported across several languages and a UK-inclusive four-country study, but by translation rather than exhaustive"}
  • {"subgroup": "longitudinal / over time", "grade": "Low", "note": "Contested: strict temporal invariance in some large samples but not established in a partial-hospital sample, complicating change-score use (Ong 2021)"}
  • {"subgroup": "occupation / industry", "grade": "Absent", "note": "No study testing invariance across occupational groups located in the latest review pass"}

Responsiveness and MIC Moderatethinevidence form: canonical

indirect responsiveness evidence comes from a chronic-depression trial and Czech psychotherapy samples, not from UK or workplace populations, and MIC anchors vary by setting

Responsiveness (sensitivity to change) is demonstrated, but formal minimal important change estimates are sparse. In a multisite chronic-depression trial (N=261), GAD-7 scores fell significantly in patients who improved on the Hamilton depression rating (effect size -0.51 at 12 weeks, -1.0 at 48 weeks) and rose in those who worsened, supporting sensitivity to change (Toussaint 2020). The Czech multi-sample study found clear sensitivity to change during psychotherapy and derived a reliable-change threshold of about plus or minus 5.2 points (Cigler 2026). A single, well-anchored minimal important change value validated for a working-adult population was not located in the latest review pass; responsiveness is established while the MIC anchor remains setting-dependent.

Populations, languages and norms

GAD-7 has been fielded and psychometrically evaluated in a very wide range of populations and languages: the 2025 Cochrane review alone drew on 27 countries and 24 languages (Akturk 2025, doi:10.1002/14651858.CD015455). General-population normative data exist, including German norms by sex and age where roughly 5% scored 10 or higher (Lowe 2008, doi:10.1097/MLR.0b013e318160d093). Validations span primary care, adolescents and older adults, perinatal women, and disease-specific groups (HIV, epilepsy, coronary heart disease, traumatic brain injury), plus students and employees. UK-relevant data come from nationally representative European samples and large UK treatment-seeking (IAPT) cohorts (Shevlin 2022, doi:10.1186/s12888-022-03787-5; Saunders 2023, doi:10.1186/s12888-023-04804-x; Delamain 2024, doi:10.1016/j.jad.2023.11.048). Dedicated UK working-population norms were not located in the latest review pass.

Criticisms and controversies

Recurring criticisms cluster around five points. First, dimensionality is unsettled: the one-factor model dominates but a cognitive-emotional versus somatic two-factor structure recurs in network analyses and some CFAs (Moriana 2021, doi:10.1002/jclp.23217; Pheh 2023, doi:10.1371/journal.pone.0285435). Second, discriminant validity against depression is weak in the sense that GAD-7 and PHQ-9 factors are highly correlated and co-occur, so the scale separates anxiety from depression less cleanly than from unrelated constructs (Stochl 2020, doi:10.1177/1073191120976863). Third, the 2025 Cochrane meta-analysis found only modest pooled sensitivity (0.64) at the standard cut-off of 10 for GAD, with pronounced heterogeneity, meaning the scale misses a meaningful share of cases at that threshold and performs worse for any anxiety disorder (sensitivity 0.48) (Akturk 2025, doi:10.1002/14651858.CD015455). Fourth, longitudinal measurement invariance is not guaranteed, which complicates the common practice of interpreting raw pre-post change scores (Ong 2021, doi:10.1177/10731911211035833). Fifth, the scale targets GAD specifically rather than the full anxiety-disorder spectrum, and its clinical/primary-care origin means any workplace deployment is outside its validation context. Organisational criterion validity is effectively untested.

References (24)

  1. Spitzer RL, Kroenke K, Williams JBW, Lowe B (2006). A brief measure for assessing generalized anxiety disorder: the GAD-7 https://doi.org/10.1001/archinte.166.10.1092
  2. Lowe B, Decker O, Muller S, et al. (2008). Validation and standardization of the Generalized Anxiety Disorder Screener (GAD-7) in the general population https://doi.org/10.1097/MLR.0b013e318160d093
  3. Plummer F, Manea L, Trepel D, McMillan D (2016). Screening for anxiety disorders with the GAD-7 and GAD-2: a systematic review and diagnostic metaanalysis https://doi.org/10.1016/j.genhosppsych.2015.11.005
  4. Akturk Z, Hapfelmeier A, Fomenko A, et al. (2025). Generalized Anxiety Disorder 7-item (GAD-7) and 2-item (GAD-2) scales for detecting anxiety disorders in adults https://doi.org/10.1002/14651858.CD015455
  5. White EJ, Karr JE (2025). Psychometric properties of the GAD-7 among college students: reliability, validity, factor structure, and measurement invariance https://doi.org/10.1037/tps0000382
  6. Cigler H, Patkova Dansova P, Javurkova A, et al. (2026). Validity and factor structure of the Czech GAD-7 across twenty samples and four independent translations https://doi.org/10.1016/j.janxdis.2026.103149
  7. Moriana JA, Jurado-Gonzalez FJ, Garcia-Torres F, et al. (2021). Exploring the structure of the GAD-7 scale in primary care patients with emotional disorders: a network analysis approach https://doi.org/10.1002/jclp.23217
  8. Riglea T, Wellman RJ, Sylvestre MP, et al. (2025). Factor structure and measurement invariance of the GAD-7 across time, sex, and language in young adults https://doi.org/10.1016/j.jad.2025.01.117
  9. Pheh KS, Tan CS, Lee KW, et al. (2023). Factorial structure, reliability, and construct validity of the Generalized Anxiety Disorder 7-item (GAD-7): evidence from Malaysia https://doi.org/10.1371/journal.pone.0285435
  10. Vogazianos P, Motrico E, Dominguez-Salas S, et al. (2022). Validation of the generalized anxiety disorder screener (GAD-7) in Cypriot pregnant and postpartum women https://doi.org/10.1186/s12884-022-05127-7
  11. De Man J, Absetz P, Sathish T, et al. (2021). Are the PHQ-9 and GAD-7 suitable for use in India? A psychometric analysis https://doi.org/10.3389/fpsyg.2021.676398
  12. Ong CW, Pierce BG, Klein KP, et al. (2021). Longitudinal measurement invariance of the PHQ-9 and GAD-7 https://doi.org/10.1177/10731911211035833
  13. Romano I, Ferro MA, Patte KA, et al. (2021). Measurement invariance of the GAD-7 and CESD-R-10 among adolescents in Canada https://doi.org/10.1093/jpepsy/jsab119
  14. Saunders R, Moinian D, Stott J, et al. (2023). Measurement invariance of the PHQ-9 and GAD-7 across males and females seeking treatment for common mental health disorders https://doi.org/10.1186/s12888-023-04804-x
  15. Delamain H, Buckman JEJ, Stott J, et al. (2024). Measurement invariance and differential item functioning of the PHQ-9 and GAD-7 between working age and older adults https://doi.org/10.1016/j.jad.2023.11.048
  16. Bolgeo T, Di Matteo R, Simonelli N, et al. (2023). Psychometric properties and measurement invariance of the 7-item General Anxiety Disorder scale (GAD-7) in an Italian coronary heart disease sample https://doi.org/10.1016/j.jad.2023.04.140
  17. Shevlin M, Butter S, McBride O, et al. (2022). Measurement invariance of the Patient Health Questionnaire (PHQ-9) and Generalized Anxiety Disorder (GAD-7) across four European countries https://doi.org/10.1186/s12888-022-03787-5
  18. Stochl J, Fried EI, Fritz J, et al. (2020). On dimensionality, measurement invariance, and suitability of sum scores for the PHQ-9 and the GAD-7 https://doi.org/10.1177/1073191120976863
  19. Teymoori A, Real R, Gorbunova A, et al. (2020). Measurement invariance of assessments of depression (PHQ-9) and anxiety (GAD-7) across sex, strata and linguistic backgrounds https://doi.org/10.1016/j.jad.2019.10.035
  20. Shih YC, Chou CC, Lu YJ, et al. (2022). Reliability and validity of the traditional Chinese version of the GAD-7 in Taiwanese patients with epilepsy https://doi.org/10.1016/j.jfma.2022.04.018
  21. Toussaint A, Husing P, Gumz A, et al. (2020). Sensitivity to change and minimal clinically important difference of the 7-item Generalized Anxiety Disorder Questionnaire (GAD-7) https://doi.org/10.1016/j.jad.2020.01.032
  22. Deady M, Collins DAJ, Johnston DA, et al. (2021). The impact of depression, anxiety and comorbidity on occupational outcomes https://doi.org/10.1093/occmed/kqab142
  23. Nyongesa MK, Mwangi P, Koot HM, et al. (2020). The reliability, validity and factorial structure of the Swahili version of the 7-item generalized anxiety disorder scale (GAD-7) among adults living with HIV from Kilifi, Kenya https://doi.org/10.1186/s12991-020-00312-4
  24. Basinska MA, Kwissa-Gajewska Z (2023). Psychometric properties of the Polish version of the Generalized Anxiety Disorder scale (GAD-7) in a non-clinical sample of employees https://doi.org/10.13075/ijomeh.1896.02104

Record notes

Overall confidence: GAD-7 has a deep, high-quality evidence base for internal consistency (High), diagnostic accuracy against clinical interview (High, though the best synthesis shows modest pooled sensitivity at cut-off 10 with high heterogeneity), a predominantly unidimensional structure (High but with a genuinely contested two-factor alternative), and cross-group measurement invariance by sex and age (High). That evidence was earned overwhelmingly in clinical, primary-care and non-UK samples, making it indirect for UK working adults, although UK general-population and IAPT data do exist. The honest gaps that most matter for a workplace registry: organisational criterion validity against objective work outcomes is Absent (the nearest evidence uses the derivative GAD-2 and self-reported outcomes); dedicated short-interval test-retest studies are sparse and mixed (Low), with the original ICC of 0.83 available only second-hand without interval or n; longitudinal invariance is contested, which bears directly on using the scale to track change; and a working-population minimal important change was not located in the latest review pass. Licence verified on 2026-07-12 from the body text of the steward's currently distributed GAD-7 English PDF (phqscreeners.com footer) and the LOINC mirror of the steward entry, not from the founding paper: free to use, Pfizer copyright retained, no permission required, which is a no-permission-required grant rather than a formal open or Creative Commons licence. Direct HTML fetches of phqscreeners.com returned HTTP 403, so confirmation rests on the steward's own distributed PDF footer and the LOINC mirror. Schema v0.2 recorded these honestly; the main tension was that most invariance/structure evidence is shared with the PHQ-9 in joint studies, which the maintainers have tagged canonical for the GAD-7 factor specifically where the analysis modelled the GAD-7 items separately.