Generalizability Theory Analysis of Local Curriculum Validation Instruments: Item Effects and Rater Consistency

Ruslina Irianty, Rustam A.R Selang, Grace Roselin Situmorang, Yetty Supriyati, Ilham Falani

Abstract


Validating assessment instruments for local curricula presents unique challenges due to context-specific content and reliance on expert judgment. This study applies Generalizability Theory (GT) to evaluate the reliability of a curriculum validation instrument designed for the Social Studies Local Plants program in Fakfak, Indonesia. A fully crossed GT design was used involving 26 items evaluated by 3 expert validators, yielding 78 observations. Each expert rated all items using a 4-point Likert scale. Variance components were estimated using a linear mixed-effects model with REML, implemented via the lme4 package in R. The analysis revealed that item variance accounted for 98.2% of total score variance (σ² = 291.97), indicating strong item discriminability. Rater variance (0.2%) and item × rater interaction (1.1%) were minimal, demonstrating high inter-rater consistency. The Generalizability Coefficient (G = 0.995) and Dependability Coefficient (Φ = 0.986) exceeded the thresholds for both relative and absolute decision-making. A D-study showed that high reliability (Φ ≥ 0.90) could be maintained with as few as 2 raters and 15 items. The instrument demonstrated excellent reliability and is suitable for evaluating local curriculum validity. Minimal rater-related variance suggests that future improvements should focus on item refinement rather than rater training. These findings support the broader use of GT in educational instrument validation, particularly in context-rich, expert-judged settings.

Keywords


generalizability theory; validation instrument; inter-rater reliability; local curriculum; educational measurement

Full Text:

PDF

References


Brennan, R. L. (2000a). (Mis)conceptions about generalizability theory. Educational Measurement: Issues and Practice, 19(1), 5–10.

Brennan, R. L. (2000b). Performance assessments from the perspective of generalizability theory. Applied Psychological Measurement, 24(4), 339–353. https://doi.org/10.1177/01466210022031706

Brennan, R. L. (2000c). Performance assessments from the perspective of generalizability theory. Applied Psychological Measurement, 24, 339–353.

Cetin, B., Guler, N., & Sarica, R. (2016). Using generalizability theory to examine different concept map scoring methods. Eurasian Journal of Educational Research, 66, 211–228.

Clayson, P. E., Carbine, K. A., Baldwin, S. A., Olsen, J. A., & Larson, M. J. (2021). Using generalizability theory and the ERP reliability analysis (ERA) toolbox for assessing test-retest reliability of ERP scores, part 1: Algorithms, framework, and implementation. International Journal of Psychophysiology, 166, 174–187. https://doi.org/10.1016/j.ijpsycho.2021.01.006

Crocker, L., & Algina, J. (1986). Introduction to Classical and Modern Test Theory. Harcourt Brace Javanovich College Publishers.

Dorathy, S., Amadioha, D. A., & Orluwene, D. G. W. (2021). Application of generalizability theory in the estimation of dependability of critical thinking scale for university students. Scholars Journal of Physics, Mathematics and Statistics, 8(9), 171–178. https://doi.org/10.36347/sjpms.2021.v08i09.003

Lertsakulbunlue, S., & Kantiwong, A. (2025). Evaluating the dependability of peer assessment in project-based learning for pre-clinical students: A generalizability theory approach. BMC Medical Education, 25(1), 260. https://doi.org/10.1186/s12909-025-06772-0

Peeters, M. J., Cor, M. K., Petite, S. E., & Schroeder, M. N. (2021). Validation evidence using generalizability theory for an objective structured clinical examination. Innovations in Pharmacy, 12(1), 15. https://doi.org/10.24926/iip.v12i1.2110

Polat, G., & Turhan, B. (2021). Applying generalizability theory in language testing. International Journal of Curriculum and Instruction, 13(3), 3344–3358.

Zhang, S. (2006). Investigating the relative effects of persons, items, sections, and languages on TOEIC score dependability. Language Testing, 23(3), 353–369.




DOI: https://doi.org/10.35445/alishlah.v17i4.8200

Refbacks

  • There are currently no refbacks.


Copyright (c) 2025 Ruslina Irianty, Rustam A.R Selang, Helli Ihsan, Yetty Supriyati, Ilham Falani

Al-Ishlah Jurnal Pendidikan Abstracted/Indexed by:

    

 


 

Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.