Utrecht Work Engagement Scale, 9-item (UWES-9)
Identity
Version: UWES-9 (9-item short form of the original 17-item UWES). An ultra-short 3-item form (UWES-3) also exists.
Structure: 9 items, scored 0 to 6 (never to always/daily); three subscales of 3 items each (vigour, dedication, absorption), commonly summed to a single engagement score.
Original citation: Schaufeli WB, Bakker AB, Salanova M (2006). The Measurement of Work Engagement With a Short Questionnaire: A Cross-National Study. Educational and Psychological Measurement. https://doi.org/10.1177/0013164405282471 (short form derived from the 17-item UWES, Schaufeli et al. 2002, https://doi.org/10.1023/a:1015630930326)
Steward / publisher: Wilmar Schaufeli and colleagues (Occupational Health Psychology Unit, Utrecht University); distributed via the author's website with an accompanying test manual.
Constructs claimed
The UWES-9 claims to measure work engagement, defined as a positive, fulfilling, work-related state of mind comprising three dimensions: vigour (energy and mental resilience while working), dedication (a sense of significance, enthusiasm and pride) and absorption (being fully concentrated and happily engrossed in work) (Schaufeli et al. 2006; Schaufeli et al. 2002). Engagement is theorised as conceptually distinct from, and in part the positive antipode of, burnout, and is embedded in the job demands-resources tradition (Crawford et al. 2010).
Evidence
Structural validity Moderatecontestedevidence form: canonical
indirectness see findings
The dimensional structure of the UWES-9 is genuinely unresolved in the literature, and this is the instrument's central psychometric controversy. The developers reported that confirmatory factor analysis supported the intended three-factor structure (vigour, dedication, absorption) across ten countries (N = 14,521), while noting the three subscales are very highly intercorrelated (Schaufeli et al. 2006). A dedicated review of 21 CFA studies of the UWES found no consensus: the three-factor structure was judged superior in 6 studies, a single general factor in a further 6, the one- and three-factor solutions were treated as equivalent in 8 studies, and 1 study confirmed neither, leading the author to warn that this ambiguity may challenge the three-factor conception of engagement itself (Kulikowski 2017). Primary studies since then remain split. A three-factor (or second-order) solution fit best in Norwegian occupational groups (Nerstad et al. 2009), Portuguese rescue workers (Sinval et al. 2018), a Vietnamese nurse sample (three-factor marginally better than one-factor, Tran et al. 2020) and Spanish health workers (three correlated factors with correlated errors, Dominguez-Salas et al. 2022). A one-factor solution was preferred in a Serbian sample (Petrovic et al. 2017) and in a German oncology-rehabilitation sample where a single factor explained 67% of variance (Sautier et al. 2015), and unidimensionality was supported using ordinal methods in Scandinavian haemodialysis nurses (Lindberg et al. 2025). Most starkly, a Swedish multi-occupational female sample (N = 702) obtained poor fit for one-, two- and three-factor models alike (RMSEA never below 0.166), i.e. no acceptable structure at all (Willmer et al. 2019). The recurring pattern is that inter-factor correlations are so high that the three subscales are difficult to separate empirically, which is why many authors recommend using the total score as a single engagement index (Mills et al. 2011).
Confidence note: Moderate. Many studies with large total N, but findings are directly contradictory on the key question (1-factor vs 3-factor), so the evidence supports 'well studied but unresolved' rather than a settled structure.
Convergent and discriminant validity Moderatecontestedevidence form: canonical
indirectness see findings
Convergent validity is reasonably supported: UWES scores correlate positively with job satisfaction, organisational commitment, meaning of work and perceived job resources, and negatively with burnout and turnover intention (Sautier et al. 2015; Hallberg & Schaufeli 2006). Discriminant validity is more contested. On the positive side, one CFA study concluded that work engagement, job involvement and organisational commitment are empirically distinct constructs reflecting different aspects of work attachment (Hallberg & Schaufeli 2006), and a meta-analytic review reported that engagement shows discriminant validity from, and incremental criterion validity over, established job attitudes (Christian et al. 2011). On the negative side, the sharpest challenge concerns overlap with burnout: a meta-analysis of 50 samples found that dimension-level correlations between burnout and engagement are high, that the two show a similar pattern of associations with correlates, and that controlling for burnout substantially reduced the effect sizes attributed to engagement, casting doubt on their functional distinctiveness (Cole et al. 2012). The instrument's own developers frame engagement partly as the positive antipode of burnout, with a best-fitting two-factor burnout-engagement model in the original cross-national data (Schaufeli et al. 2006), which is consistent with, rather than resolving, the overlap concern. Discriminant validity against related well-being constructs (workaholism, job boredom) was more clearly demonstrated for the ultra-short UWES-3 across five national samples (Schaufeli et al. 2019).
Confidence note: Moderate. Convergent evidence is consistent and includes meta-analysis; discriminant evidence is genuinely mixed, with a credible meta-analytic challenge (burnout overlap) that is not fully answered.
Criterion validity: reference standard Lowthinevidence form: canonical
indirectness see findings
Criterion validity against hard organisational and health outcomes is thinner than the volume of UWES research implies, and most evidence is cross-sectional or predictive of self-reported rather than registered outcomes. The strongest workplace-relevant evidence against a registered health outcome comes from a one-year prospective cohort of 4,921 employees: baseline UWES scores were negatively associated with register-recorded long-term sickness absence due to mental illness, but discrimination was only moderate (area under the ROC curve = 0.70) and below the pre-set threshold for practical screening use (0.75); crucially, UWES scores were NOT associated with sickness absence due to musculoskeletal or other somatic illness, so predictive validity was specific to mental-illness absence (Roelen et al. 2014). For turnover, engagement measured with the UWES is repeatedly associated with lower turnover intention, but through cross-sectional or mediational designs rather than actual turnover: examples include Chinese nurses during COVID-19 (Tang et al. 2022) and surgical trainees' intention to leave training (Dominguez et al. 2018). At meta-analytic level, engagement relates positively to task and contextual performance and mediates demands/resources effects on performance, though these syntheses pool multiple engagement measures rather than the UWES-9 alone (Christian et al. 2011). A UK quality-improvement evaluation found a modest but statistically significant difference in UWES scores between intervention and control wards (White et al. 2014). Evidence linking UWES scores prospectively to diagnosed conditions, objective productivity or actual (not intended) turnover is sparse.
Confidence note: Low. One good prospective registry study exists (with an informative null for somatic absence); most other criterion evidence is cross-sectional, uses intentions rather than behaviour, or pools multiple engagement measures.
Criterion validity: organisational Lowthinevidence form: canonical
indirectness see findings
Organisational criterion evidence (sickness absence, turnover, performance, diagnosed conditions in a work context): see the criterion findings; graded from the pass-one record.
Confidence note: Low. One good prospective registry study exists (with an informative null for somatic absence); most other criterion evidence is cross-sectional, uses intentions rather than behaviour, or pools multiple engagement measures.
Internal consistency Highwell-establishedevidence form: canonical
indirectness see findings
Internal consistency is the UWES-9's most consistently strong property. The developers reported the three subscale scores had good internal consistency across the cross-national samples (Schaufeli et al. 2006). Total-score Cronbach's alpha is typically in the low-to-mid 0.90s: 0.93 (ordinal alpha) in Scandinavian haemodialysis nurses (Lindberg et al. 2025), 0.93 for the total scale in Vietnamese nurses (subscales 0.86 vigour, 0.77 absorption, 0.90 dedication) (Tran et al. 2020), and 0.94 in the German oncology-rehabilitation sample (Sautier et al. 2015). Reliability was likewise satisfactory in Norwegian (Nerstad et al. 2009) and Serbian (Petrovic et al. 2017) validations. A caveat: alpha values in the 0.90s for a 9-item scale with very high inter-item correlation partly reflect redundancy, and high total-scale alpha does not adjudicate the one- versus three-factor question. Where the three-item subscales are used separately, the absorption subscale tends to be the weakest (Tran et al. 2020).
Confidence note: High. Multiple good-quality studies across many languages and occupations, large total N, consistently reporting total-score alpha in the low-to-mid 0.90s.
Test-retest reliability Lowthin
indirectness see summary
Genuine test-retest reliability evidence for the UWES-9 is comparatively scarce, and this is the weakest-documented property relative to the instrument's popularity. The developers' original short-form paper states summarily that the three UWES-9 scores have 'good' test-retest reliability across the cross-national dataset, but reports no coefficients or retest intervals in the material available for this record (Schaufeli et al. 2006). Among UWES-9 validations, one of the few to report a short-interval coefficient found an intraclass correlation of only 0.48 over roughly three months in a Vietnamese nurse subsample, which is modest and below conventional stability thresholds (Tran et al. 2020). A multisample, longitudinal construct-validity study of the UWES exists and, by its title, addresses longitudinal/temporal evidence (Seppala et al. 2009); however, its full text and abstract could not be retrieved in this pass, so no specific stability coefficient from it is asserted here, and any stability figures it reports pertain principally to the 17-item UWES rather than the 9-item form. Many national validation studies report internal consistency and factor structure but omit test-retest data entirely. Engagement is theorised as a relatively stable state, so a stronger, replicated test-retest evidence base would be expected for the 9-item form than currently exists.
Confidence note: Low. Test-retest evidence specific to the UWES-9 is sparse: the developer claim is summary and coefficient-free in the material accessed, the clearest located UWES-9 retest coefficient (ICC 0.48 over ~3 months) is modest, and longer-interval stability evidence is tied to the 17-item form and was not retrievable in the latest review pass. The thinness of a replicated UWES-9 test-retest base is itself the finding.
Measurement invariance Moderatewell-establishedevidence form: canonical
indirectness see findings
Measurement invariance has been tested repeatedly, most often across country/language and occupational group, with configural and metric invariance commonly achieved and scalar invariance less consistently so. Full-scale measurement invariance for the UWES-9 was obtained between Portuguese and Brazilian workers for both the three-factor first-order and second-order models, permitting direct mean comparisons (Sinval et al. 2018). In Portuguese rescue workers, the UWES-9 first-order model reached full (uniqueness) invariance across occupational groups, while the second-order model reached only partial (metric) second-order invariance (Sinval et al. 2018). Gender invariance has been supported in Spanish health-care professionals (Dominguez-Salas et al. 2022), and factorial invariance across ten occupational groups was reported for the Norwegian UWES (Nerstad et al. 2009). The developers' cross-national work established the UWES-9 across ten countries but treated cross-national equivalence descriptively rather than through the full modern invariance hierarchy (Schaufeli et al. 2006). Against this, a Rasch analysis of the 17-item UWES in Korea found that only 9 of 17 items differentiated adequately between men and women, i.e. differential item functioning by sex (Song et al. 2020), a caution that invariance is not universal. Over-time (longitudinal) invariance is less frequently formally tested; the multisample longitudinal study is the main source addressing temporal stability of the structure (Seppala et al. 2009).
Confidence note: Moderate. Several studies reach metric and sometimes scalar/full invariance across language and occupation, but scalar invariance is not universal, some DIF by sex is reported, and formal over-time invariance testing is limited.
Responsiveness and MIC Very lowthinevidence form: canonical
indirectness see findings
Formal responsiveness (sensitivity to change) and a minimal important change (MIC) value have not been established for the UWES-9. No located study derived an anchor-based or distribution-based MIC, and the instrument was not developed as an outcome measure with defined change metrics. Indirect evidence of sensitivity to change comes from a UK evaluation of the 'Productive Ward' quality-improvement programme, where UWES scores were modestly but significantly higher in intervention wards than matched controls (4.33 vs 4.07, p = 0.013), which the authors interpreted as the UWES being able to detect programme-related differences (White et al. 2014). This is a between-group cross-sectional contrast rather than a within-person responsiveness or MIC analysis. No UK or other MIC benchmark was located in this pass.
Confidence note: Very low. No MIC established and no formal responsiveness study located; only indirect, between-group evidence that scores can differ with an intervention.
Populations, languages and norms
The UWES-9 has been validated in a very wide range of languages and occupations, giving broad but heterogeneous population coverage. The developers' short-form study drew on samples from ten countries (N = 14,521) (Schaufeli et al. 2006). Located validations span Norwegian occupational groups (Nerstad et al. 2009), Serbian employees (Petrovic et al. 2017), Brazilian and Portuguese workers (Sinval et al. 2018), Portuguese rescue workers (Sinval et al. 2018), Vietnamese hospital nurses (Tran et al. 2020), Spanish health-care professionals (Dominguez-Salas et al. 2022), a German oncology-rehabilitation patient sample (Sautier et al. 2015), a Swedish multi-occupational female sample (Willmer et al. 2019) and Scandinavian (Danish and Swedish) haemodialysis nurses (Lindberg et al. 2025); an ultra-short UWES-3 was validated across Finland, Japan, the Netherlands, Belgium/Flanders and Spain (Schaufeli et al. 2019) and in Peru (Merino-Soto et al. 2022). UK-specific evidence is limited: the clearest UK-context deployment located is the English-language 'Productive Ward' nursing study (conducted in Ireland with a UK/Ireland health-service context), which used the UWES and reported mean scores around 4.1 to 4.3 on the 0 to 6 metric, but this is a study sample, not a representative UK norm (White et al. 2014). No representative UK normative dataset was located in this pass; the reference norm tables that exist are those in the authors' international test manual (UWES Test Manual, Schaufeli & Bakker 2003), which are international rather than UK-specific.
Criticisms and controversies
Three recurring criticisms appear in the literature. First, factor-structure instability: a review of 21 CFA studies found no consensus on whether the UWES is one-dimensional or three-dimensional, with roughly equal support for each and one study confirming neither, which the author argued could undermine the three-dimensional conception of engagement itself (Kulikowski 2017); an extreme case obtained no acceptable fit for any of the one-, two- or three-factor models (Willmer et al. 2019). Second, discriminant validity from burnout: a meta-analysis of 50 samples reported high dimension-level correlations, near-identical correlate patterns, and shrinking engagement effect sizes once burnout is controlled, questioning whether engagement and burnout are functionally distinct (Cole et al. 2012). The developers' own framing of engagement as the positive antipode of burnout, with a best-fitting combined two-factor model, sits uneasily with claims of full independence (Schaufeli et al. 2006). Third, redundancy and subscale separability: because the three subscales correlate so highly, several authors conclude the total score should be used and that the subscales add little, so reporting 'vigour/dedication/absorption' as three distinct measures may over-claim (Mills et al. 2011). A related psychometric caution is that Cronbach's alpha in the 0.90s for such highly intercorrelated items partly reflects item redundancy rather than only reliability. Discriminant concerns against neighbouring attitudes (job involvement, organisational commitment) are, by contrast, comparatively reassuring (Hallberg & Schaufeli 2006).
References (27)
- Schaufeli WB, Bakker AB, Salanova M (2006). The Measurement of Work Engagement With a Short Questionnaire: A Cross-National Study https://doi.org/10.1177/0013164405282471
- Schaufeli WB, Salanova M, González-Romá V, Bakker AB (2002). The Measurement of Engagement and Burnout: A Two Sample Confirmatory Factor Analytic Approach https://doi.org/10.1023/a:1015630930326
- Seppälä P, Mauno S, Feldt T, Hakanen J, Kinnunen U, Tolvanen A, Schaufeli W (2009). The Construct Validity of the Utrecht Work Engagement Scale: Multisample and Longitudinal Evidence https://doi.org/10.1007/s10902-008-9100-y
- Schaufeli WB, Shimazu A, Hakanen J, Salanova M, De Witte H (2019). An Ultra-Short Measure for Work Engagement: The UWES-3 https://doi.org/10.1027/1015-5759/a000430
- Kulikowski K (2017). Do we all agree on how to measure work engagement? Factorial validity of Utrecht Work Engagement Scale as a standard measurement tool: A literature review https://doi.org/10.13075/ijomeh.1896.00947
- Nerstad CGL, Richardsen AM, Martinussen M (2009). Factorial validity of the Utrecht Work Engagement Scale (UWES) across occupational groups in Norway https://doi.org/10.1111/j.1467-9450.2009.00770.x
- Willmer M, Westerberg Jacobson J, Lindberg M (2019). Exploratory and Confirmatory Factor Analysis of the 9-Item Utrecht Work Engagement Scale in a Multi-Occupational Female Sample https://doi.org/10.3389/fpsyg.2019.02771
- Lindberg M, Knudsen K, Lindberg M (2025). Factor structure of the Utrecht Work Engagement Scale in a sample of Danish and Swedish haemodialysis nurses https://doi.org/10.1186/s12912-025-03545-4
- Sinval J, Marques-Pinto A, Queirós C, Marôco J (2018). Work Engagement among Rescue Workers: Psychometric Properties of the Portuguese UWES https://doi.org/10.3389/fpsyg.2017.02229
- Petrović IB, Vukelić M, Čizmić S (2017). Work Engagement in Serbia: Psychometric Properties of the Serbian Version of the Utrecht Work Engagement Scale (UWES) https://doi.org/10.3389/fpsyg.2017.01799
- Sinval J, Pasian S, Queirós C, Marôco J (2018). Brazil-Portugal Transcultural Adaptation of the UWES-9: Internal Consistency, Dimensionality, and Measurement Invariance https://doi.org/10.3389/fpsyg.2018.00353
- Mills MJ, Culbertson SS, Fullagar CJ (2011). Conceptualizing and Measuring Engagement: An Analysis of the Utrecht Work Engagement Scale https://doi.org/10.1007/s10902-011-9277-3
- Sautier LP, Scherwath A, Weis J, Sarkar S, Bosbach M, Schendel M, Ladehoff N, Koch U, Mehnert A (2015). Assessment of Work Engagement in Patients with Hematological Malignancies: Psychometric Properties of the German Version of the UWES-9 https://doi.org/10.1055/s-0035-1555912
- Song HD, Hong AJ, Jo Y (2020). Psychometric Investigation of the Utrecht Work Engagement Scale-17 Using the Rasch Measurement Model https://doi.org/10.1177/0033294120922494
- Hallberg UE, Schaufeli WB (2006). "Same Same" But Different? Can Work Engagement Be Discriminated from Job Involvement and Organizational Commitment? https://doi.org/10.1027/1016-9040.11.2.119
- Cole MS, Walter F, Bedeian AG, O'Boyle EH (2012). Job Burnout and Employee Engagement: A Meta-Analytic Examination of Construct Proliferation https://doi.org/10.1177/0149206311415252
- Christian MS, Garza AS, Slaughter JE (2011). Work Engagement: A Quantitative Review and Test of Its Relations with Task and Contextual Performance https://doi.org/10.1111/j.1744-6570.2010.01203.x
- Crawford ER, LePine JA, Rich BL (2010). Linking job demands and resources to employee engagement and burnout: A theoretical extension and meta-analytic test https://doi.org/10.1037/a0019364
- Roelen CAM, van Hoffen MFA, Groothoff JW, de Bruin J, Schaufeli WB, van Rhenen W (2014). Can the Maslach Burnout Inventory and Utrecht Work Engagement Scale be used to screen for risk of long-term sickness absence? https://doi.org/10.1007/s00420-014-0981-2
- White M, Wells JS, Butterworth T (2014). The impact of a large-scale quality improvement programme on work engagement: preliminary results from a national cross-sectional survey of the 'Productive Ward' https://doi.org/10.1016/j.ijnurstu.2014.05.002
- Tang Y, Dias Martins LM, Wang SB, He QX, Huang HH (2022). The impact of nurses' sense of security on turnover intention during the normalization of COVID-19 epidemic: The mediating role of work engagement https://doi.org/10.3389/fpubh.2022.1051895
- Dominguez LC, Stassen L, de Grave W, Sanabria A, Alfonso E, Dolmans D (2018). Taking control: Is job crafting related to the intention to leave surgical training? https://doi.org/10.1371/journal.pone.0197276
- Tran TTT, Watanabe K, Imamura K, Nguyen HT, Sasaki N, Kuribayashi K, Sakuraya A, et al. (2020). Reliability and validity of the Vietnamese version of the 9-item Utrecht Work Engagement Scale https://doi.org/10.1002/1348-9585.12157
- Domínguez-Salas S, Rodríguez-Domínguez C, Arcos-Romero AI, Allande-Cussó R, et al. (2022). Psychometric Properties of the Utrecht Work Engagement Scale (UWES-9) in a Sample of Active Health Care Professionals in Spain https://doi.org/10.2147/PRBM.S387242
- Merino-Soto C, Lozano-Huamán M, Lima-Mendoza S, Calderón de la Cruz G, Juárez-García A, Toledano-Toledano F (2022). Ultrashort Version of the Utrecht Work Engagement Scale (UWES-3): A Psychometric Assessment https://doi.org/10.3390/ijerph19020890
- Schaufeli WB, Desart S, De Witte H (2020). Burnout Assessment Tool (BAT): Development, Validity, and Reliability https://doi.org/10.3390/ijerph17249495
- Schaufeli WB, Bakker AB (2003). UWES Utrecht Work Engagement Scale: Preliminary Manual (Version 1) https://www.wilmarschaufeli.nl/publications/Schaufeli/Test%20Manuals/Test_manual_UWES_English.pdf
Record notes
[Upgraded from v0.1 to v0.2 structure in pass two; criterion field split, licence re-verified 2026-07-12.] Overall confidence: internal consistency is High; structural validity, convergent/discriminant validity and measurement invariance are Moderate but each carries a genuine, cited contradiction rather than clean support; criterion validity is Low; test-retest is Low with the absence itself a finding; responsiveness/MIC is Very low (effectively absent). Points the schema made hard to record honestly: (1) The one-factor versus three-factor debate is not a defect to be resolved to a single 'confidence' but a standing feature of the evidence; the 'structural_validity' field forces a single grade onto directly contradictory findings, so 'Moderate' here means 'extensively studied but unresolved', not 'moderately good'. (2) 'constructs_claimed' presents three subscales, yet a substantial strand of evidence argues the subscales are not empirically separable and only the total should be used; the field cannot easily flag that the claimed structure is itself contested. (3) Several strong stability and invariance findings pertain to the 17-item UWES, not the 9-item form; the schema does not distinguish evidence transferred from the parent instrument from evidence earned by the UWES-9 directly, and the maintainers have flagged this inline where it applies. (4) Test-retest and MIC are near-absent for the 9-item form specifically despite the instrument's popularity, which is a more important finding than a coefficient would have been. (5) UK-specific norms were not located as a representative benchmark; the 'Productive Ward' study is UK/Ireland health-service context but is a study sample, not a norm, and the only reference norms are the authors' international (non-UK) test manual. (6) Per the registry rule on 'clinical': the UWES-9 has been deployed in patient-rehabilitation samples (for example haematological malignancy patients), but it is a work-engagement instrument, not a clinical screener, and none of its validation was earned in a diagnostic setting; workplace deployment should be read as a non-clinical context throughout.