multi-item-scale · reviewed 2026-07-12 contested: Structural validity

General Health Questionnaire-12 (GHQ-12)

Licence verified: 2026-07-12 · record reviewed: 2026-07-12 (pass two)

Identity

Version: GHQ-12 (12-item short form of Goldberg's General Health Questionnaire; the family comprises GHQ-60, GHQ-30, GHQ-28 and GHQ-12)

Structure: 12

Original citation: Goldberg DP (1972) The Detection of Psychiatric Illness by Questionnaire, Maudsley Monograph No. 21, Oxford University Press; the 12-item short form was later derived from this work. Validity of the two short versions summarised by Goldberg et al. 1997.

Steward / publisher: GL Assessment (part of the GL Education Group), United Kingdom, copyright holder; translated versions distributed internationally by the Mapi Research Trust via ePROVIDE. Part of the permissions payment is paid as a royalty to the Institute of Psychiatry.

Licence status (verified 2026-07-12): Proprietary and commercial (licensed, paid). The GHQ is not open access or public domain. The steward's current copyright statement reads that the General Health Questionnaire is protected worldwide by international copyright laws in all languages, with all rights reserved to GL Assessment, UK, and must not be used without permission; permission for licensees in the United Kingdom, Republic of Ireland and Channel Islands is obtained from GL Assessment (permissions@gl-assessment.co.uk) and, for other countries, from the Mapi Research Trust via ePROVIDE (https://eprovide.mapi-trust.org). GL Assessment's current permissions policy states that all GL Assessment products are protected by copyright and may not usually be reproduced in hard copy or electronic form, or translated, without permission. GL Assessment's product page markets the GHQ range (GHQ-12, GHQ-28, GHQ-30, GHQ-60) as a purchasable assessment requiring professional-eligibility registration to order. The steward's GHQ support FAQ states that photocopying a record form without abiding by the permission conditions is regarded as theft and a criminal offence, that part of the permissions payment is paid as a royalty to the Institute of Psychiatry, and that translated versions are distributed by the Mapi Research Trust but have not been validated by GL Assessment. Verified at the last verification attempt against the steward's current pages, not founding papers.
Source: GL Assessment GHQ product page (https://www.gl-assessment.co.uk/products/general-health-questionnaire/, page current May 2026); GL Assessment permissions policy (https://www.gl-assessment.co.uk/policies/permissions/); GL Assessment / GL Education GHQ support FAQ page (https://support.gl-assessment.co.uk/knowledge-base/assessments/general-health-questionnaire-support/about-the-general-health-questionnaire/faqs), which carries the 'photocopying a record form is regarded as theft and a criminal offence', Institute of Psychiatry royalty, and Mapi-distributed-unvalidated-translations statements; and the GL Assessment copyright/permission notice ('all rights reserved to GL Assessment, UK; do not use without permission; UK/Ireland/Channel Islands contact permissions@gl-assessment.co.uk, other countries contact the Mapi Research Trust'). All read in the latest review pass via web search of the steward's live pages.

Constructs claimed

Non-specific psychiatric morbidity / common mental disorder, that is current psychological distress. The GHQ detects breaks in normal healthy functioning and the appearance of new distressing symptoms over the recent past (roughly the last few weeks), rather than long-standing or enduring conditions. Marketed and used as a screener for minor psychiatric disorder in community, primary-care and occupational-health settings.

Evidence

Deployment context caveat. The GHQ-12 is a psychiatric-morbidity screener developed and validated for case detection in general medical, primary-care and community settings; its origin and validation are stated here as fact, but workplace deployment is a different context and any occupational use must be treated as such. It measures short-term states (symptoms of less than roughly two weeks are captured, and the instrument does not detect chronic or long-standing conditions), so it is a distress screener rather than a diagnostic or wellbeing measure. It is also a proprietary, licensed instrument requiring paid permission, which contrasts with the free instruments in this registry and constrains organisational deployment (cost, permission, and prohibition on uncontrolled reproduction). (applies to every property below)

Structural validity Highcontestedevidence form: canonical

direct - large representative general-population and survey samples (England, Spain, Germany) plus meta-analytic aggregation; the definitive evidence is not workplace-specific but the structural conclusion is not population-dependent.

The internal structure is the central and long-running controversy for this instrument, and it is driven by scoring method rather than by any true multidimensionality. Although Goldberg designed the GHQ-12 as a unidimensional measure, factor-analytic studies have reported one-, two- and three-factor solutions, the best known being Graetz's three factors (anxiety/depression, social dysfunction, loss of confidence) and various two-factor solutions splitting positively and negatively worded items. Hankins 2008 showed in Health Survey for England data (n=3705) that the best fitting model is one-dimensional with a response-bias factor on the negatively phrased items, concluding that earlier multifactor structures were artefacts of the analysis method and of the mixing of positively and negatively worded items across scoring schemes (Likert, GHQ 0-0-1-1 and C-GHQ). Rey et al. 2014 replicated this in a large Spanish sample (n=27,674), finding that spurious multidimensionality appears only under corrected and Likert scoring because of ambiguous response categories in the negative items, and recommending standard GHQ scoring with a single global score. Romppel et al. 2013 reached the same conclusion in a German population sample (N=2041), with subscale correlations against external criteria (BDI, PHQ-2, SF-36) not differing substantially from one another. The strongest evidence comes from Gnambs and Staufenbiel 2018, two meta-analyses (summary data from 38 studies, total N=76,473; and individual responses from 84 samples, N=410,640): confirmatory and bifactor modelling showed that although two wording-based factors are recoverable, almost all common variance loads on a general factor, so the GHQ-12 is essentially unidimensional and subscale scores should not be interpreted. Hystad and Johnsen 2020 corroborated the bifactor structure (general factor plus positive- and negative-wording method factors) in military samples. The modern consensus is therefore essential unidimensionality with method effects from item wording; the historical multifactor literature is treated here as contested and largely superseded.

Convergent and discriminant validity Moderatewell-establishedevidence form: canonical

indirect - convergent coefficients come from non-UK, non-workplace samples (India, Australia) and clinical cancer cohorts; no UK working-adult convergent study was located in the latest review pass.

Convergent evidence is consistent though less systematically studied than structure or criterion validity. The GHQ-12 total correlates moderately with other distress and wellbeing measures: r=0.58 with the Subjective Well-being Inventory in older Indian adults (Qin et al. 2018). In head-to-head comparisons against diagnostic interviews the GHQ-12 discriminates common mental disorders comparably to the SRQ, K10, K6 and PHQ (Patel et al. 2007; Gill et al. 2007), although Gill et al. 2007 found the SF-12 mental component and the K6/K10 outperformed the GHQ-12 for detecting diagnosed depression. In UK cancer patients the GHQ-12 tracked the Distress Thermometer and HADS over time (Gessler et al. 2008). Discriminant separation from distinct constructs is less directly documented in the sources retrieved.

Criterion validity: reference standard Highwell-establishedevidence form: canonical

indirect - reference-standard validation is extensive but drawn from general medical, primary-care, community and international samples, and (for Baksheev) adolescents; no criterion validation against a diagnostic standard in a UK working-adult sample was located in the latest review pass.

Validity against a diagnostic reference standard is the GHQ-12's best-evidenced property. In the WHO multi-centre study of mental illness in general health care (5438 patients, 15 centres, CIDI primary-care interview), the GHQ-12 achieved a mean area under the ROC curve of 0.88 (range 0.83 to 0.95), performing as well as the longer GHQ-28, with complex scoring offering no advantage and no significant effect of gender, age or education on validity (Goldberg et al. 1997). Against the CIS-R in Indian primary care the GHQ-12 was among the best of five screeners, though positive predictive value at optimal cut-offs was modest across all five screeners compared, ranging from 51% to 77% (Patel et al. 2007). Against the CIDI in an Australian general-population survey (N=10,504) the GHQ-12 discriminated depression and anxiety, albeit less well than the K6/K10 for depression (Gill et al. 2007). In adolescents against a SCID DSM-IV-TR interview the AUC was 0.781, with sex-specific optimal thresholds (Baksheev et al. 2011). Case-detection performance is thus robust and reproducible across languages and settings, but is anchored in primary-care, community and clinical samples rather than workplaces.

Criterion validity: organisational Absent (a finding about the literature)untestedevidence form: canonical

direct - the absence applies squarely to the UK working-adult deployment question; no evidence located to grade.

No study validating the GHQ-12 against organisational outcomes (sickness absence, staff turnover, job performance, or diagnosed conditions recorded in a work context) as a criterion was located in the latest review pass. The GHQ-12 is very widely used within occupational and workforce samples, for example among South African healthcare workers where its reliability and dimensionality were examined (Kufe et al. 2024) and in UK primary-care detection studies (Plummer et al. 2000), but in those studies the GHQ is deployed as the distress measure itself, not validated against a separate work-outcome standard. The absence of organisational criterion evidence is reported here as the finding it is: the instrument's predictive relationship to absence, turnover or performance in the workplace was not established in the retrieved literature.

Internal consistency Highwell-establishedevidence form: canonical

indirect - alpha estimates come from England (general population), India and international student/community samples rather than UK working adults specifically, though the property is stable across populations.

Internal consistency is consistently high. Cronbach's alpha of about 0.90 has been reported under Likert and GHQ scoring in Health Survey for England data, falling to 0.75 under C-GHQ scoring (Hankins 2008); alpha of 0.9 in older Indian adults (Qin et al. 2018); and a lower KR-20 of 0.70 under dichotomous scoring in Colombian students (Simancas-Pallares et al. 2017). Values across the wider literature typically sit in the 0.80 to 0.90 range. An important caveat is that Hankins 2008 showed Cronbach's alpha overestimates reliability once the negative-item response bias is modelled, so the true measurement precision is somewhat lower than the headline alpha values suggest.

Test-retest reliability Lowthin

indirect - the single located estimate is from Japanese young adults over a two-week interval, not UK working adults; it is also a two-way random-effects ICC for agreement rather than a workplace test-retest.

CoefficientTypeIntervalSamplePopulationEvidence form
greater than 0.70ICC2 weeks137 analysed (154 tested at baseline)young adults, Japan (non-clinical)canonical

Genuine test-retest (stability) evidence for the GHQ-12 is thin. Only one clean estimate was located in the latest review pass: an intraclass correlation exceeding 0.70 over a two-week interval in 137 young adults, reported alongside a standard error of measurement of 1.47 (bimodal scoring) and 2.44 (Likert scoring) and corresponding smallest detectable change values (Ohno et al. 2017). Because the GHQ measures a current, changeable state rather than a trait, high test-retest coefficients are not necessarily expected, and the scarcity of stability data is itself a graded finding. Reported alpha values from other studies must not be mistaken for retest reliability.

Measurement invariance Lowthinevidence form: mixed

indirect - invariance evidence comes from military, international and cross-language samples; occupation-level and UK working-adult invariance were not established in the latest review pass.

Invariance evidence is partial. Goldberg et al. 1997 found no significant effect of gender, age or educational level on the validity of the GHQ across the WHO multi-centre sample, and validity coefficients held across ten translated languages, supporting broad cross-population comparability of case detection. Hystad and Johnsen 2020 demonstrated that a bifactor structure was invariant across two independent military samples and, in a multi-group analysis, across time. The meta-analytic confirmation of a common structure across 84 samples (Gnambs and Staufenbiel 2018) is consistent with structural stability across populations. However, formal configural/metric/scalar invariance testing across sex, age band and, in particular, occupation is not comprehensively established in the retrieved literature, and translated versions are noted by the steward as not all validated by the publisher.

Sub-grades (evidence differs by subgroup):

  • {"subgroup": "sex / age / education", "grade": "Moderate", "note": "no effect on validity across the WHO multi-centre study (Goldberg et al. 1997); reasonably supported."}
  • {"subgroup": "across time / repeated samples", "grade": "Low", "note": "bifactor structure invariant across two military samples and over time (Hystad and Johnsen 2020); limited to one military cohort."}
  • {"subgroup": "occupation", "grade": "Absent", "note": "no formal invariance test across occupational groups located in the latest review pass."}

Responsiveness and MIC Very lowthinevidence form: canonical

indirect - the change evidence comes from a UK cancer cohort and Japanese young adults, not UK working adults; no anchor-based minimal important change exists.

Formal responsiveness and a minimal important change value are not established. Gessler et al. 2008 found that GHQ-12 scores changed over four and eight weeks in the same direction as the HADS and Distress Thermometer in UK cancer outpatients, providing indirect support for sensitivity to change. Ohno et al. 2017 derived smallest detectable change values from the standard error of measurement (for example a smallest detectable change of about 4.06 points at the individual level under bimodal scoring), which bounds interpretable change but is a distribution-based statistic, not an anchor-based minimal important change. No validated minimal important change for the GHQ-12 was located in the latest review pass.

Populations, languages and norms

The GHQ-12 is one of the most widely translated and used distress screeners internationally, with validated or examined versions across, among many others, English general-population and survey samples (Hankins 2008), German (Romppel et al. 2013), Spanish (Rey et al. 2014), Chinese adolescents (Li et al. 2009), Ukrainian refugees (Benoni et al. 2024), older Indian adults (Qin et al. 2018) and South African healthcare workers (Kufe et al. 2024). UK normative data are provided in the publisher's GHQ user guide (a commercial document, not verified in the latest review pass); the steward notes that translated versions are distributed by the Mapi Research Trust but have not all been validated by GL Assessment. Optimal cut-off scores vary by population, scoring method and sex, so a single universal threshold should not be assumed.

Criticisms and controversies

Three related criticisms dominate. First, the scoring-method controversy: the choice between binary GHQ (0-0-1-1), Likert (0-1-2-3) and C-GHQ scoring materially changes the apparent factor structure and reliability, and much of the historical multi-factor literature is now regarded as an artefact of scoring and of the mix of positively and negatively worded items (Hankins 2008; Rey et al. 2014; Gnambs and Staufenbiel 2018). Practitioners should pre-specify scoring and treat the GHQ-12 as a single total score, not interpret subscales. Second, response bias on the negatively phrased items inflates Cronbach's alpha and can create spurious dimensions, so reported reliabilities overstate true precision (Hankins 2008). Third, as a screener the GHQ-12 has modest positive predictive value at realistic prevalence (the five screeners compared, including the GHQ, gave PPVs of 51 to 77% in primary care) and detects only recent-onset distress, so it is not a diagnostic instrument and misses long-standing conditions (Patel et al. 2007). For this registry's audience two further points matter: it is a licensed commercial product (unlike the free instruments here), and its criterion evidence is clinical/primary-care in origin with no validation against organisational outcomes.

References (18)

  1. Gnambs T; Staufenbiel T (2018). The structure of the General Health Questionnaire (GHQ-12): two meta-analytic factor analyses https://doi.org/10.1080/17437199.2018.1426484
  2. Hankins M (2008). The reliability of the twelve-item general health questionnaire (GHQ-12) under realistic assumptions https://doi.org/10.1186/1471-2458-8-355
  3. Rey JJ; Abad FJ; Barrada JR; Garrido LE; Ponsoda V (2014). The impact of ambiguous response categories on the factor structure of the GHQ-12 https://doi.org/10.1037/a0036468
  4. Romppel M; Braehler E; Roth M; Glaesmer H (2013). What is the General Health Questionnaire-12 assessing? Dimensionality and psychometric properties of the GHQ-12 in a large scale German population sample https://doi.org/10.1016/j.comppsych.2012.10.010
  5. Goldberg DP; Gater R; Sartorius N; Ustun TB; Piccinelli M; Gureje O; Rutter C (1997). The validity of two versions of the GHQ in the WHO study of mental illness in general health care https://doi.org/10.1017/s0033291796004242
  6. Hystad SW; Johnsen BH (2020). The Dimensionality of the 12-Item General Health Questionnaire (GHQ-12): Comparisons of Factor Structures and Invariance Across Samples and Time https://doi.org/10.3389/fpsyg.2020.01300
  7. Baksheev GN; Robinson J; Cosgrave EM; Baker K; Yung AR (2011). Validity of the 12-item General Health Questionnaire (GHQ-12) in detecting depressive and anxiety disorders among high school students https://doi.org/10.1016/j.psychres.2010.10.010
  8. Simancas-Pallares MA; Arrieta Vergara KM; Arevalo Tovar L (2017). Construct validity and internal consistency of three factor structures and two scoring methods of the 12-item General Health Questionnaire https://doi.org/10.7705/biomedica.v37i3.3240
  9. Qin T; Vlachantoni A; Evandrou M; Falkingham J (2018). General Health Questionnaire-12 reliability, factor structure, and external validity among older adults in India https://doi.org/10.4103/psychiatry.IndianJPsychiatry_112_17
  10. Patel V; Araya R; Chowdhary N; King M; Kirkwood B; Nayak S; Simon G; Weiss HA (2007). Detecting common mental disorders in primary care in India: a comparison of five screening questionnaires https://doi.org/10.1017/S0033291707002334
  11. Gill SC; Butterworth P; Rodgers B; Mackinnon A (2007). Validity of the mental health component scale of the 12-item Short-Form Health Survey (MCS-12) as measure of common mental disorders in the general population https://doi.org/10.1016/j.psychres.2006.11.005
  12. Gao W; Stark D; Bennett MI; Seymour J; Higginson IJ (2011). Using the 12-item General Health Questionnaire to screen psychological distress from survivorship to end-of-life care: dimensionality and item quality https://doi.org/10.1002/pon.1989
  13. Ohno S; Takahashi K; Inoue A; Takada K; Ishihara Y; Tanigawa M; Hirao K (2017). Smallest detectable change and test-retest reliability of a self-reported outcome measure: results of the CES-D, General Self-Efficacy Scale, and 12-item General Health Questionnaire https://doi.org/10.1111/jep.12795
  14. Plummer S; Gournay K; Goldberg D; Ritter S; Mann A; Blizard R (2000). Detection of psychological distress by practice nurses in general practice https://doi.org/10.1017/s0033291799002597
  15. Gessler S; Low J; Daniells E; Williams R; Brough V; Tookman A; Jones L (2008). Screening for distress in cancer patients: is the distress thermometer a valid measure in the UK and does it measure change over time? A prospective validation study https://doi.org/10.1002/pon.1273
  16. Benoni R; Sartorello A; Mazzi M; Berti G; et al. (2024). The use of 12-item General Health Questionnaire (GHQ-12) in Ukrainian refugees: translation and validation study of the Ukrainian version https://doi.org/10.1186/s12955-024-02226-1
  17. Li WHC; Chung JOK; Chui MML; Chan PSL (2009). Factorial structure of the Chinese version of the 12-item General Health Questionnaire in adolescents https://doi.org/10.1111/j.1365-2702.2009.02905.x
  18. Kufe CN; Bernstein K; Wilson KS (2024). Reliability, validity and dimensionality of the 12-Item General Health Questionnaire among South African healthcare workers https://doi.org/10.4102/ajopa.v6i0.144

Record notes

Overall confidence: the GHQ-12 is among the best-evidenced distress screeners for internal structure (High, though contested historically) and criterion validity against diagnostic interviews (High), with high internal consistency (High). It is materially weaker on the properties this registry cares about for a UK working-adult audience: test-retest stability rests on a single small non-UK study (Low/thin), measurement invariance across occupation is untested, responsiveness and minimal important change are not formally established (Very low), and organisational criterion validity against work outcomes is Absent. All grades other than structural validity are downgraded to indirect because the evidence is overwhelmingly non-UK, non-workplace, or clinical/primary-care in origin, and the instrument's clinical-screener heritage and short-term state focus are flagged in the deployment caveat. Licence verified in the latest review pass against GL Assessment and GL Education current pages and the Mapi Research Trust listing: it is a proprietary, paid, permission-required product, an important contrast to the free tools in the registry. Schema v0.2 made honest recording straightforward here; the main friction was that UK working-adult norms and the definitive user-guide psychometrics sit behind a paywalled commercial manual that could not be inspected in the latest review pass, so those are flagged rather than asserted.