Psychometric evaluation of the Arabic language fluency assessment (ALFA): Evidence from classical test theory

Authors

  • Dagi Mardhan Agae Agni UIN Sunan Ampel Surabaya
  • M. Baihaqi

DOI:

https://doi.org/10.30603/al.v11i2.7830

Keywords:

Arabic language testing;, CEFR;, classical test theory;, language assessment;, psychometric analysis

Abstract

Background: Language assessment is widely used to determine learners’ proficiency; however, the accuracy of score interpretation depends on the quality of the measurement instrument.

Aims: This study aims to evaluate the psychometric quality of the Arabic Language Fluency Assessment (ALFA) using Classical Test Theory. The psychometric findings are subsequently interpreted from the perspective of communicative language assessment informed by the CEFR framework.

Methods: A quantitative descriptive-evaluative design was employed using Classical Test Theory (CTT). Data were collected from 92 beginner-level learners and analyzed using item difficulty, discrimination index, reliability (Cronbach’s Alpha), and distractor functioning.

Results: The findings show that the test is dominated by easy items, with mean difficulty indices of 0.93 and 0.88, resulting in limited score variability. Consistent with this pattern, most items exhibited poor discrimination, while a large proportion of distractors were non-functioning. Although the reliability coefficients were relatively high (0.885 and 0.843), these values indicate internal consistency rather than overall measurement quality.

Implications: Within the present sample, the ALFA demonstrated limited ability to differentiate learner performance. The psychometric findings also raise the possibility of construct underrepresentation when interpreted from the perspective of communicative language assessment informed by the CEFR framework. The findings support systematic revision of the ALFA instrument and may also inform the improvement of institution-developed Arabic language assessments in comparable educational contexts.

Downloads

Download data is not yet available.

References

Alamer, A., Al Khateeb, A., & Alshabeb, A. (2025). The creation and validation of the Arabic vocabulary levels Test (Arabic-VLT). Language Assessment Quarterly, 22(2), 117–137. https://doi.org/10.1080/15434303.2025.2477445

AlKhuzaey, S., Grasso, F., Payne, T. R., & Tamma, V. (2024). Text-based question difficulty prediction: A systematic review of automatic approaches. International Journal of Artificial Intelligence in Education, 34(3), 862–914. https://doi.org/10.1007/s40593-023-00362-1

Al-Owidha, A. A. (2018). Investigating the psychometric properties of the qiyas for L1 Arabic language test using a rasch measurement framework. Language Testing in Asia, 8(1), 12. https://doi.org/10.1186/s40468-018-0064-5

Ayanwale, M. A., Chere-Masopha, J., & Morena, M. C. (2022). The classical test or item response measurement theory: The status of the framework at the examination council of Lesotho. International Journal of Learning, Teaching and Educational Research, 21(8), 384–406. https://doi.org/10.26803/ijlter.21.8.22

Bahnson, M., Sallai, G., Jwa, K., & Berdanier, C. G. P. (2023). Mitigating ceiling effects in a longitudinal study of doctoral engineering student stress and persistence. International Journal of Doctoral Studies, 18, 199–227. https://doi.org/10.28945/5118

Bond, T. G., & Fox, C. M. (2021). Applying the rasch model: Fundamental measurement in the human sciences. Routledge. https://doi.org/10.4324/9781410614575

Brown, H. D., & Abeywickrama, P. (2019). Language assessment: Principles and classroom practices. Pearson.

Cardwell, R., Naismith, B., LaFlair, G. T., & Nydick, S. (2023). Duolingo English test: Technical manual. Duolingo Research Report, 1–33.

Chapelle, C. A. (2021). Argument-based validation in testing and assessment. SAGE Publications. https://doi.org/10.4135/9781071878811

Council of Europe. (2020). Common european framework of reference for languages: Learning, teaching, assessment. Council of Europe Publishing.

Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555

DeVellis, R. F. (2017). Scale development: Theory and applications. SAGE Publications.

Devianti, R., Subiyanto, D., & Ratna Purnamarini, T. (2025). Digital transformation as a driver of innovative behavior: The mediating roles of hr analytics and psychological well-being at the library and archives office of Bantul regency. International Journal of Economics and Management Review, 3(3), 46–60. https://doi.org/10.58765/ijemr.v3i3.275

Ebel, R. L., & Frisbie, D. A. (1991). Essentials of educational measurement (5th ed.). Prentice Hall.

El Chaal, R., Seghir, R. Ben, & Aboutafail, M. O. (2025). Psychometric evaluation of human-crafted and ai-generated multiple-choice questions for Mathematics instruction. Education Science and Management, 3(2), 78–92. https://doi.org/10.56578/esm030201

Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists. Lawrence Erlbaum Associates.

Fulcher, G. (2010). Practical language testing. Hodder Education. https://doi.org/10.4324/9780203767399

Fulcher, G., & Harding, L. (2022). The Routledge handbook of language testing (2nd ed.). Routledge. https://doi.org/10.1080/13803611.2023.2179073

Girolamo, T., Ghali, S., Campos, I., & Ford, A. (2022). Interpretation and use of standardized language assessments for diverse school-age individuals. Perspectives of the ASHA Special Interest Groups, 7(4), 981–994. https://doi.org/10.1044/2022_PERSP-21-00322

Haladyna, T. M., & Rodriguez, M. C. (2013). Developing and validating test items. Routledge. https://doi.org/10.4324/9780203850381

Kremmel, B., & Harding, L. (2020). Towards a comprehensive, empirical model of language assessment literacy across stakeholder groups: Developing the language assessment literacy survey. Language Assessment Quarterly, 17(1), 100–120. https://doi.org/10.1080/15434303.2019.1674855

Kunnan, A. J. (2020). Evaluating language assessments from an ethics perspective. JLTA Journal, 23, 3–13. https://doi.org/10.20622/jltajournal.23.0_3

Ljubojević, D., & Daničić, M. (2026). Assessing oral proficiency in young EFL learners: A study of eighth-grade pupils’ performance at the serbian national English competition. Research in Pedagogy, 16(1), 22–35. https://doi.org/10.5937/IstrPed2601022L

Ludewig, U., Schwerter, J., & McElvany, N. (2023). The features of plausible but incorrect options: Distractor plausibility in synonym-based vocabulary tests. Journal of Psychoeducational Assessment, 41(7). https://doi.org/10.17877/DE290R-24398

Najmalia Fitra, Herdah, H., & Mahrous, A. E. (2025). Language test item analysis techniques: Teknik analisis item tes bahasa. Al Mahāra: Jurnal Pendidikan Bahasa Arab, 11(1), 158–179. https://doi.org/10.14421/almahara.2025.0111-09

Nasr, I. (2026). Balancing formative and summative assessment in IB MYP Arabic B: A qualitative case study of teachers’ practices and perspectives. An-Najah University Journal for Research – B (Humanities). https://doi.org/10.35552/0247.a2830

Rakhlin, N. V., Aljughaiman, A., & Grigorenko, E. L. (2021). Assessing language development in Arabic: The Arabic language: Evaluation of Function (ALEF). Applied Neuropsychology: Child, 10(1), 37–52. https://doi.org/10.1080/21622965.2019.1596113

Rezigalla, A. A., Eleragi, A. M. E. S. A., Elhussein, A. B., Alfaifi, J., ALGhamdi, M. A., Al Ameer, A. Y., Yahia, A. I. O., Mohammed, O. A., & Adam, M. I. E. (2024). Item analysis: The impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items. BMC Medical Education, 24(1), 1–7. https://doi.org/10.1186/s12909-024-05433-y

Rodriguez, M. C. (2005). Three options are optimal for multiple-choice items: A meta-analysis of 80 years of research. Educational Measurement: Issues and Practice, 24(2), 3–13. https://doi.org/10.1111/j.1745-3992.2005.00006.x

Schmucker, R., & Moore, S. (2026). The impact of item-writing flaws on difficulty and discrimination in item response theory. Computers and Education: Artificial Intelligence, 11, 100632. https://doi.org/10.1016/j.caeai.2026.100632

Sharma, L. R. (2021). Analysis of difficulty index, discrimination index and distractor efficiency of multiple-choice questions of speech sounds of English. International Research Journal of MMC, 2(1), 15–28. https://doi.org/10.3126/irjmmc.v2i1.35126

Sireci, S. G., & Rodriguez, G. (2022). Validity in educational testing (pp. 1–9). https://doi.org/10.4324/9781138609877-ree180-1

Subando, J. (2025). A multidimensional item response theory approach in the item analysis of Arabic language tests in Madrasah Aliyah. Jurnal Penelitian dan Evaluasi Pendidikan, 29(2), 271–284. https://doi.org/10.21831/pep.v29i2.90877

Yu, G., & Xu, J. (Eds.). (2024). Language test validation in a digital age. Cambridge University Press.

Zeinoun, P., Iliescu, D., & El Hakim, R. (2022). Psychological tests in Arabic: A review of methodological practices and recommendations for future use. Neuropsychology Review, 32(1), 1–19. https://doi.org/10.1007/s11065-021-09476-6

Zhao, H., & Aryadoust, V. (2025). A meta-analysis of the reliability of second language reading comprehension assessment tools. Studies in Second Language Acquisition, 47(1), 388–416. https://doi.org/10.1017/S0272263124000627

Downloads

Published

2026-08-26

How to Cite

Agae Agni, D. M., & M. Baihaqi. (2026). Psychometric evaluation of the Arabic language fluency assessment (ALFA): Evidence from classical test theory. Al-Lisan: Jurnal Bahasa (e-Journal), 11(2), 237–250. https://doi.org/10.30603/al.v11i2.7830