Reliability generalization of scores on the L2 grit scale: A meta-analytic investigation
Research Methods in Applied Linguistics, cilt.5, sa.3, 2026 (ESCI, Scopus)
- Yayın Türü: Makale / Derleme
- Cilt numarası: 5 Sayı: 3
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.rmal.2026.100347
- Dergi Adı: Research Methods in Applied Linguistics
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus
- Anahtar Kelimeler: Cronbach’s alpha, Internal consistency, L2 grit, Meta-analysis, Reliability generalization
- Yıldız Teknik Üniversitesi Adresli: Evet
Özet
This study examines the internal consistency of scores obtained from the L2 Grit Scale (L2GS; Teimouri, Plonsky, et al., 2022) through a reliability generalization (RG) meta-analysis. A systematic search of Web of Science and Scopus (2020–2026) identified 76 eligible studies comprising 81 independent datasets. A random-effects model was applied to 59 Cronbach's alpha coefficients after 24 coefficient-level observations were excluded due to non-comparable, incomplete, or non-extractable reliability reporting. The findings yielded a weighted mean alpha of α = .834 (95% CI [.820, .847]), indicating acceptable mean internal consistency across applications. The 95% prediction interval [.738, .930] — representing the expected reliability range for a new, independent application of the scale — is wider and more practically informative than the confidence interval, underscoring that reliability cannot be assumed uniform across contexts. The mean alpha is broadly consistent with benchmarks reported in the SLA literature, although such values are better interpreted as context-sensitive heuristics than fixed evaluative criteria (McNeish, 2018; Plonsky & Derrick, 2016). Heterogeneity was substantial (Q (58) = 1043.09, p ' .001; I ² = 95.27%; τ = .047), indicating considerable variability in reliability estimates across studies. Potential selective-reporting effects were examined via funnel-plot inspection, Egger's regression test, and trim-and-fill; none suggested substantial asymmetry, though such diagnostics should be interpreted cautiously in the RG context. A sensitivity analysis comparing untransformed and Bonett-transformed coefficients confirmed the robustness of the primary findings. Overall, the substantial heterogeneity underscores the need for context-sensitive interpretation of L2GS reliability estimates.