COPSOQ III (Copenhagen Psychosocial Questionnaire, third version, core and dimension scales)
Identity
Version: COPSOQ III (international core, middle and long versions; national validated versions derived from these), published 2019 by the international COPSOQ network
Structure: Version-dependent. The international COPSOQ III is organised as a set of dimension scales in three tiers (core, middle, long). National validated versions vary: the German middle version has 84 items across 31 scales (Lincke 2021); the Greek long version has 108 items across 40 scales (Kotsakis 2025). Most dimensions are short multi-item scales (typically two to four items); several constructs are measured with a single item.
Original citation: Burr H, Berthelsen H, Moncada S, Nübling M, et al. The Third Version of the Copenhagen Psychosocial Questionnaire. Safety and Health at Work 2019 (Burr 2019)
Steward / publisher: COPSOQ International Network (coordinated internationally; national COPSOQ teams steward validated language versions, e.g. FFAW Freiburg for the German version)
Constructs claimed
Psychosocial working conditions (exposures), organised as many largely independent dimension scales grouped into domains: demands at work (quantitative, work pace, emotional, demands for hiding emotions); work organisation and job content (influence, possibilities for development, variation, meaning of work, commitment to the workplace); interpersonal relations and leadership (predictability, recognition, role clarity, role conflicts, quality of leadership, social support, sense of community); work-individual interface (job insecurity, insecurity over working conditions, job satisfaction, work-life/work-privacy conflict); social capital (vertical and horizontal trust, organisational justice); offensive behaviours (bullying, sexual harassment, threats, violence); plus outcome/strain scales (self-rated health, burnout, stress, sleeping troubles, and, in longer versions, work engagement). COPSOQ III is an exposure measure of the psychosocial work environment rather than a wellbeing outcome instrument, although it embeds a small number of health and strain outcome scales.
Evidence
Structural validity Moderatewell-establishedevidence form: mixed
indirect : evidence is from Turkish, Australian, Norwegian, Swedish and (for COPSOQ-II) Australian and Polish samples; no UK-specific structural validation was retrieved, and some of the strongest structural evidence is parent-form (COPSOQ-II).
COPSOQ III is built as a set of largely independent dimension scales rather than one higher-order factor, so structural validity is assessed dimension by dimension, and the network explicitly notes that factor structure for the newly introduced dimensions was still to be fully tested at launch (Burr 2019). Confirmatory and exploratory factor analyses of national versions generally support the intended per-dimension structure. The Turkish COPSOQ-3 reported an excellent-fitting model (RMSEA 0.038, SRMR 0.053, CFI 0.98) with 19 extracted factors explaining 66.1% of variance (Sahan 2018). The Australian long-version validation, using EFA followed by CFA, reported acceptable to good fit (a four-factor higher-order EFA solution, and CFA RMSEA around 0.034 to 0.036 with CFI/TLI in the 0.89 to 0.92 range depending on model) (Rahimi 2025). For the predecessor COPSOQ-II, a set-ESEM approach that permitted only theory-consistent cross-loadings improved fit markedly over strict CFA (from CFI 0.907 to CFI 0.947 to 0.971) and reduced inflated inter-factor correlations, indicating that strict independent-clusters CFA understates fit for this multidimensional instrument (Dicke 2018). A full CFA of the COPSOQ-II 33-subscale model in a Polish sample also showed good fit (RMSEA below 0.05, SRMR below 0.08) (Baka 2022). The Norwegian COPSOQ III validation combined CFA with item response theory to characterise dimension structure and item information (Ose 2023). The main structural caveat is that fit is evaluated separately per dimension or per national selection of dimensions, not for a single global model, and that some short dimensions are less stable across countries.
Sub-grades (evidence differs by subgroup):
- {"subgroup": "Established core dimensions (demands, influence, leadership, social support, social capital)", "grade": "Moderate", "note": "Consistently recovered as intended factors across national CFAs (Sahan 2018; Rahimi 2025; Baka 2022)."}
- {"subgroup": "Newly introduced COPSOQ III dimensions (e.g. work engagement, quality of work, cyber bullying)", "grade": "Low", "note": "Network flagged their factor structure as not yet fully tested at launch (Burr 2019); evidence still accumulating."}
Convergent and discriminant validity Moderatewell-establishedevidence form: mixed
indirect : convergent evidence against ERI is from a German general-population employed sample ([Nuebling 2013](https://doi.org/10.1186/1471-2458-13-538)); the strongest discriminant analysis is parent-form (COPSOQ-II) ([Dicke 2018](https://doi.org/10.3389/fpsyg.2018.00584)); no UK sample.
Convergent and discriminant validity are supported mainly through scale intercorrelations and comparison with other established work-stress models. In the Gutenberg Health Study, COPSOQ scales and the Effort-Reward Imbalance (ERI) questionnaire showed congruent patterns across occupational groups, and COPSOQ predictor scales explained comparable or slightly greater variance than ERI in shared outcome scales (for example job satisfaction R-squared 0.51 for COPSOQ versus 0.46 for ERI; burnout 0.35 versus 0.26), supporting convergent validity against an established alternative model (Nuebling 2013). Within-instrument, developers calculate scale intercorrelations to demonstrate that dimensions are distinct (divergent) while theoretically related scales correlate (convergent) (Burr 2019; Berthelsen 2020). The set-ESEM analysis of COPSOQ-II is directly relevant to discriminant validity: strict CFA produced some inflated inter-factor correlations that fell substantially once theory-consistent cross-loadings were permitted, indicating the dimensions are discriminable but not perfectly orthogonal (Dicke 2018). A dedicated study constructed and validated a global Workplace Social Capital scale from COPSOQ III justice and trust items, supporting convergent structure among the social-capital dimensions (Berthelsen 2019). Formal multitrait-multimethod or correlations with independent external gold-standard constructs are sparser than the within-instrument evidence.
Criterion validity: reference standard Not applicable (category difference)untestedevidence form: canonical
direct : the category (diagnostic reference-standard validity) does not apply to a psychosocial-exposure measure; recorded as Not-applicable rather than Absent because it is a category mismatch, and no reference-standard study was located.
COPSOQ III is an exposure measure of psychosocial working conditions, not a screener for a health condition, so validity against a diagnostic or clinical reference standard is not a design goal and is largely not evaluated. No study retrieved in the latest review pass compared COPSOQ III dimension scores against a diagnostic reference standard (for example a structured clinical interview for depression) with sensitivity or specificity. The embedded strain scales (self-rated health, burnout, stress) are self-report and are treated as outcomes correlated with exposures rather than validated against a clinical criterion. This is an appropriate absence for an exposure instrument, but it means the field is genuinely untested rather than merely weak.
Criterion validity: organisational Very lowthinevidence form: mixed
indirect : the only quantified associations are cross-sectional and against self-reported outcomes in a German general-population sample ([Nuebling 2013](https://doi.org/10.1186/1471-2458-13-538)); objective-outcome prediction is unstudied and no UK data exist.
Evidence that COPSOQ III scores predict objective organisational or health outcomes (register-based sickness absence, staff turnover, diagnosed conditions, performance) is a recognised gap that the developers themselves flag as future work. The Swedish validation explicitly states that evaluating predictive criterion validity against register data on absence, staff turnover and performance remains to be done in longitudinal multilevel designs (Berthelsen 2020), and the Norwegian study likewise lists predictive validity as not yet established for their version (Ose 2023). The available criterion-type evidence is cross-sectional and against self-reported outcomes: in the Gutenberg Health Study COPSOQ psychosocial scales explained meaningful variance in self-reported job satisfaction (R-squared 0.51), burnout (0.35), satisfaction with life (0.18) and general health (0.11), with theoretically sensible predictors (for example meaning of work and sense of community for job satisfaction; work-privacy conflict for burnout) (Nuebling 2013). This supports concurrent association with self-reported strain outcomes but not prediction of objective work outcomes. Genuine organisational criterion validity against registers or hard outcomes is essentially Absent in the retrieved literature.
Internal consistency Highwell-establishedevidence form: mixed
indirect for the UK specifically (evidence from Sweden, Germany, Australia, Turkey, Greece, China and the seven-country international sample; no UK sample), but the property is directly and repeatedly measured on the fielded COPSOQ III versions.
Internal consistency is the most extensively documented property and is generally acceptable to good for the multi-item dimensions, with well-characterised weak spots at the short two-item and emotion-related scales. Across the seven-country international middle version, most of the 23 tested scales reached Cronbach alpha above 0.70; three fell below: Commitment to the Workplace (two items, mean alpha 0.64, 95% CI 0.61 to 0.67), Demands for Hiding Emotions (three items, mean alpha 0.66, 0.58 to 0.73) and Control over Working Time, with additional country-specific shortfalls for Predictability, Meaning of Work and Job Insecurity (all two-item scales, alpha around 0.62 to 0.66 in France and Turkey) (Burr 2019). The German middle version reported that 20 of 25 multi-item scales exceeded alpha 0.70 and 13 reached 0.80 or higher; a small number of two-item scales (for example Degrees of Freedom, reduced to two items, alpha 0.53) were weak (Lincke 2021). The Australian long version found all 31 three-or-more-item scales acceptable except Demands for Hiding Emotions (0.66), and among two-item scales used Spearman-Brown coefficients, with Variation of Work unacceptably low (0.24) while Meaning of Work was acceptable (0.78) (Rahimi 2025). The Turkish version reported 23 dimensions above 0.70 with Control over Working Time (0.54) and Predictability (0.66) below (Sahan 2018). The Greek long version found 22 of 40 scales with alpha above 0.70 (Kotsakis 2025), and the Chinese long version reported an overall alpha of 0.92 with per-dimension values from 0.60 to 0.92 (Huang 2025). The consistent pattern is that longer scales are reliable and the recurring weak scales are the two-item and hiding-emotions dimensions, which is a structural feature of the compact design rather than a country-specific defect.
Sub-grades (evidence differs by subgroup):
- {"subgroup": "Multi-item dimensions (three or more items)", "grade": "High", "note": "Alpha consistently above 0.70, frequently above 0.80 (Burr 2019; Lincke 2021; Rahimi 2025)."}
- {"subgroup": "Two-item and hiding-emotions scales", "grade": "Low", "note": "Recurring alpha/Spearman-Brown shortfalls (e.g. Commitment 0.64, Hiding Emotions 0.66, Variation of Work 0.24 in Australia) (Burr 2019; Rahimi 2025)."}
Test-retest reliability Lowthin
indirect : no test-retest study of the fielded COPSOQ III was located; the adequate ICC evidence is from the Danish parent COPSOQ over about three weeks ([Thorsen 2010](https://doi.org/10.1177/1403494809349859)), and the only longitudinal COPSOQ-II retest used a 12-month interval that conflates true change with instability ([Baka 2022](https://doi.org/10.1371/journal.pone.0262266)).
| Coefficient | Type | Interval | Sample | Population | Evidence form |
|---|---|---|---|---|---|
| 0.70 to 0.89 (all scales but one adequate/good; mutual-trust scale 0.64) | ICC | median 22 days (range 6 to 65 days) | 349 respondents (283 employees) | Danish general working population (COPSOQ, parent version) | parent |
| 0.15 to 0.34 (Pearson correlations across 33 subscales; all significant at p<0.001 but low, attributed to the long interval) | r | about 12 months | 599 human-service employees | Polish human-service staff (COPSOQ II, parent version) | parent |
Test-retest reliability for COPSOQ III specifically is essentially Absent and is repeatedly named as outstanding future work. The COPSOQ III network states that test-retest of the newly introduced dimensions was still to come (Burr 2019), and the Norwegian, Greek and Chinese validations all explicitly did not assess test-retest (Ose 2023; Kotsakis 2025; Huang 2025); the German validation deliberately excluded a formal test-retest for practical reasons (Lincke 2021). The two structured findings recorded here are both parent-form: Thorsen and Bjorner's dedicated Danish test-retest study of the original COPSOQ found ICCs of 0.70 to 0.89 (one scale, mutual trust, 0.64) over a median 22-day interval, which is good (Thorsen 2010); the Polish COPSOQ-II longitudinal study reported low retest correlations (r 0.15 to 0.34) but over a 12-month interval that the authors note is far longer than the recommended few-weeks-to-months window, so it indexes real exposure change as much as instrument instability (Baka 2022). Direct short-interval COPSOQ III test-retest evidence has not been established.
Measurement invariance Very lowthinevidence form: mixed
indirect : the only explicit longitudinal invariance evidence is parent-form (COPSOQ-II, Australian principals) ([Dicke 2018](https://doi.org/10.3389/fpsyg.2018.00584)); cross-country comparability for COPSOQ III is argued by design and similarity of solutions, not by pooled invariance testing; no UK data.
Formal measurement invariance testing (configural, metric, scalar) for COPSOQ III is thin, and the international network was candid that it could not directly test differential item functioning across countries at launch because of data-protection constraints on pooling item-level data, while observing that scale properties differed somewhat across working populations (Burr 2019). The clearest invariance-relevant evidence is parent-form: the COPSOQ-II set-ESEM study of Australian school principals tested longitudinal factor structure across two time points and reported good, stable fit (for example CFI 0.95, TLI 0.94, RMSEA 0.02 at Time 1), supporting configural stability over time (Dicke 2018). Comparability across the many national COPSOQ III versions is handled largely by design (a common obligatory core, standardised translation procedures) and demonstrated indirectly through similar factor solutions and reliability patterns across countries (Burr 2019; Rahimi 2025; Kotsakis 2025) rather than by pooled multi-group invariance models. Group-level aggregation properties (ICC(1)/ICC(2) across occupations and workplaces) were examined in Sweden to justify comparing group mean scores, which is a related but distinct property from measurement invariance (Berthelsen 2020). Full multi-group scalar invariance across sex, age, occupation and language for the fielded COPSOQ III has not been established in the retrieved literature.
Sub-grades (evidence differs by subgroup):
- {"subgroup": "Longitudinal (over time)", "grade": "Low", "note": "Configural stability supported for parent COPSOQ-II via set-ESEM (Dicke 2018)."}
- {"subgroup": "Across countries/languages", "grade": "Very low", "note": "Not formally tested at item level; DIF analysis was precluded by data-protection constraints (Burr 2019)."}
- {"subgroup": "Across sex/age/occupation", "grade": "Absent", "note": "No formal multi-group invariance study located for COPSOQ III."}
Responsiveness and MIC Absent (a finding about the literature)untestedevidence form: canonical
indirect : the only relevant material is a proposed 5 to 10 point minimum important difference convention from the Swedish benchmarking work ([Berthelsen 2020](https://doi.org/10.3390/ijerph17093179)); no formal responsiveness study and no UK data.
Responsiveness (sensitivity to change over time or after intervention) and a minimal important change threshold have not been established for COPSOQ III in the retrieved literature. The Swedish validation lists responsiveness alongside test-retest and predictive validity as properties still to be evaluated in future longitudinal designs (Berthelsen 2020). As an interpretation aid rather than a formal responsiveness statistic, the Swedish work proposes that a 5 to 10 point difference on the 0 to 100 scale metric can be treated as a minimum important difference, and reports Cohen's d effect sizes for group contrasts (Berthelsen 2020), but this is a benchmarking convention, not an anchor-based or distribution-based minimal-important-change study. No study retrieved in the latest review pass estimated responsiveness against an external change anchor.
Populations, languages and norms
COPSOQ III has been validated and normed in a wide range of national and language versions, and benchmark/reference values are a core feature of the system. Retrieved validations include Swedish (national benchmarks) (Berthelsen 2020), German (database exceeding 250,000 participants) (Lincke 2021), Norwegian (registered nurses) (Ose 2023), Portuguese (municipal and healthcare workers; and a 2026 national validation) (Cotrim 2022; Cotrim 2026), Turkish (Sahan 2018), Greek (Kotsakis 2025), Australian (national benchmarks by ANZSCO group) (Rahimi 2025), Chinese (Huang 2025) and Czech (Zabrodska 2026) versions, building on the earlier German COPSOQ database tradition (Nuebling 2010). Occupation-specific and country-specific reference values exist for several versions. Critically for this audience, no UK-specific COPSOQ III validation or UK norm set was located in the latest review pass; the nearest English-language reference data are Australian (Rahimi 2025).
Criticisms and controversies
Several recurring critiques emerge from the retrieved literature. First, COPSOQ III is a modular family of dimension scales rather than a single validated scale, so psychometric quality is heterogeneous across dimensions; the compact two-item and hiding-emotions scales repeatedly show sub-0.70 internal consistency (Burr 2019; Rahimi 2025), and reliability of two-item scales must be read via Spearman-Brown rather than alpha. Second, key longitudinal properties are missing: test-retest reliability, responsiveness and predictive criterion validity against objective outcomes are explicitly named as not yet done by the developers themselves (Berthelsen 2020; Ose 2023), so the instrument's status as an exposure measure rests largely on cross-sectional and construct-validity evidence. Third, measurement invariance across countries could not be tested at item level at launch owing to data-protection constraints on pooling data (Burr 2019), leaving cross-national comparability argued by design rather than demonstrated. Fourth, strict independent-clusters CFA tends to understate fit and inflate factor correlations for this instrument, which has prompted use of ESEM/set-ESEM approaches (Dicke 2018). Fifth, for a UK working-adult audience the evidence base is population-indirect: there is no retrieved UK validation, and much of the strongest reliability and invariance evidence is parent-form (COPSOQ or COPSOQ-II) rather than the fielded COPSOQ III.
References (18)
- Burr H, Berthelsen H, Moncada S, Nübling M, Dupret E, Demiral Y, Oudyk J, Kristensen TS, Llorens C, Navarro A, Lincke HJ, Bocéréan C, Sahan C, Smith P, Pohrt A, International COPSOQ Network (2019). The Third Version of the Copenhagen Psychosocial Questionnaire https://doi.org/10.1016/j.shaw.2019.10.002
- Berthelsen H, Westerlund H, Bergström G, Burr H (2020). Validation of the Copenhagen Psychosocial Questionnaire Version III and Establishment of Benchmarks for Psychosocial Risk Management in Sweden https://doi.org/10.3390/ijerph17093179
- Lincke HJ, Vomstein M, Lindner A, Nolle I, Häberle N, Haug A, Nübling M (2021). COPSOQ III in Germany: validation of a standard instrument to measure psychosocial factors at work https://doi.org/10.1186/s12995-021-00331-1
- Ose SO, Lohmann-Lafrenz S, Bernström VH, Berthelsen H, Marchand GH (2023). The Norwegian version of the Copenhagen Psychosocial Questionnaire (COPSOQ III): Initial validation study using a national sample of registered nurses https://doi.org/10.1371/journal.pone.0289739
- Cotrim TP, Bem-Haja P, Pereira A, Fernandes C, Azevedo R, Fonte C, Nossa P, Silva CF, Barbosa-Ferreira J, Silvério J (2022). The Portuguese Third Version of the Copenhagen Psychosocial Questionnaire: Preliminary Validation Studies of the Middle Version among Municipal and Healthcare Workers https://doi.org/10.3390/ijerph19031167
- Rahimi E, Arnold KA, LaMontagne AD, et al. (2025). Validation and benchmarks for the Copenhagen Psychosocial Questionnaire (COPSOQ III) in an Australian working population sample https://doi.org/10.1186/s12889-025-21845-x
- Kotsakis R, Avraam E, Malliarou M, et al. (2025). A Validation Study of the COPSOQ III Greek Questionnaire for Assessing Psychosocial Factors in the Workplace https://doi.org/10.3390/healthcare13161980
- Şahan C, Baydur H, Demiral Y (2018). A novel version of Copenhagen Psychosocial Questionnaire-3: Turkish validation study https://doi.org/10.1080/19338244.2018.1538095
- Huang Y, Zhang Y, Wang X, et al. (2025). COPSOQ III in China: Preliminary Validation of an International Instrument to Measure Psychosocial Work Factors https://doi.org/10.3390/healthcare13070825
- Thorsen SV, Bjorner JB (2010). Reliability of the Copenhagen Psychosocial Questionnaire https://doi.org/10.1177/1403494809349859
- Dicke T, Marsh HW, Riley P, Parker PD, Guo J, Horwood M (2018). Validating the Copenhagen Psychosocial Questionnaire (COPSOQ-II) Using Set-ESEM: Identifying Psychosocial Risk Factors in a Sample of School Principals https://doi.org/10.3389/fpsyg.2018.00584
- Nübling M, Seidler A, Garthus-Niegel S, Latza U, Wagner M, Hegewald J, Liebers F, Jankowiak S, Zwiener I, Wild PS, Letzel S (2013). The Gutenberg Health Study: measuring psychosocial factors at work and predicting health and work-related outcomes with the ERI and the COPSOQ questionnaire https://doi.org/10.1186/1471-2458-13-538
- Baka Ł, Prusik M, Pejtersen JH (2022). Full evaluation of the psychometric properties of COPSOQ II. One-year longitudinal study on Polish human service staff https://doi.org/10.1371/journal.pone.0262266
- Berthelsen H, Westerlund H, Pejtersen JH, Hadzibajramovic E (2019). Construct validity of a global scale for Workplace Social Capital based on COPSOQ III https://doi.org/10.1371/journal.pone.0221893
- Nübling M, Hasselhorn HM (2010). The Copenhagen Psychosocial Questionnaire in Germany: from the validation of the instrument to a national survey https://doi.org/10.1177/1403494809353652
- Kristensen TS, Hannerz H, Høgh A, Borg V (2005). The Copenhagen Psychosocial Questionnaire, a tool for the assessment and improvement of the psychosocial work environment https://doi.org/10.5271/sjweh.948
- Cotrim TP, Bem-Haja P, Vagos P, et al. (2026). Validation of the third version of the Copenhagen Psychosocial Questionnaire for Portugal https://doi.org/10.1371/journal.pgph.0006036
- Zábrodská K, Květon P, Jelínek M, et al. (2026). Psychometric validation of the Czech Copenhagen Psychosocial Questionnaire (COPSOQ III) https://doi.org/10.1186/s40359-026-03961-4
Record notes
Overall confidence: Moderate for structural and convergent/discriminant validity, High for internal consistency and for the breadth of populations/norms, but Very low to Absent for test-retest, measurement invariance, responsiveness/MIC and organisational criterion validity. COPSOQ III should be recorded as a well-established psychosocial-exposure instrument whose cross-sectional and internal-consistency credentials are strong but whose longitudinal and predictive properties remain genuinely under-evidenced, a gap the developers themselves acknowledge. The record is deliberately framed like an exposure measure (comparable to the HSE Management Standards Indicator Tool) rather than a wellbeing outcome. Two honesty points the v0.2 schema still made awkward. (1) The instrument is a family of dimension scales, so most single record-level grades summarise a heterogeneous per-dimension reality; subgrades capture the most important splits (multi-item versus two-item scales; established versus newly introduced dimensions), but a fully faithful record would grade each dimension separately, which the single-record structure cannot hold. (2) Much of the reliability and invariance evidence is parent-form (original COPSOQ or COPSOQ-II); evidence_form is tagged 'mixed' or 'parent' throughout so this is not read as canonical COPSOQ III evidence, and the recurring theme is that COPSOQ III inherits confidence from its lineage while its own longitudinal evidence base is still thin. Licence: the questionnaire is released under Creative Commons CC BY-NC-ND 4.0, verified 2026-07-12 by fetching and reading the full text of the COPSOQ International Network's current 'Licence, Guidelines & Questionnaire' page and the guidelines PDF linked from it, not from founding papers. The lead's 'free for non-commercial with registration' pointer was an oversimplification in a different direction than a simple registration model: the licence is CC BY-NC-ND (non-commercial, no-derivatives) but the network's own page explicitly permits commercial use provided no fee is charged for the questionnaire itself (fees for assessment, advice, analysis and training are allowed), requires prior contact with the national COPSOQ team before a new translation or adaptation, and prohibits distributing modified material as 'COPSOQ'. The guidelines PDF adds a carve-out that the Work Engagement items (WE_T, WE1 to WE3) may be used commercially only under separate agreement with Triple i (3ihc.nl), because they derive from the Schaufeli et al. Utrecht Work Engagement short scale. Population indirectness: no UK COPSOQ III validation was located; the closest English-language validation and norms are Australian (Rahimi 2025). All grades are flagged indirect accordingly.