ChatGPT for Grading Students’ Writing: A Systematic Review of Evidence in Relation to Core Language Assessment Principles

نویسندگان

1 Universitas Pendidikan Indonesia

2 Universitas Pendidikan Indonesia

3 Universitas Pendidikan Indonesia

4 Universitas Pendidikan Indonesia

5 Universitas Pendidikan Indonesia

6 Universitas Pendidikan Indonesia

doi
10.22034/ijlt.2026.571744.1510
چکیده

Writing assessment has developed immensely, particularly in the AI era, where the evaluation of students’ writing can be conducted with greater efficiency and speed. As a generative AI language model, ChatGPT has increasingly been used for writing assessment purposes, particularly due to its efficiency, immediacy, and scalability, which may support self-revision and self-regulated learning. Nevertheless, the quality of its feedback and grading in relation to the six core principles of language assessment, namely validity, reliability, fairness, washback, authenticity, and practicality, remains contested. This systematic literature review synthesizes evidence from 24 Scopus-indexed ELT studies published between 2023 and 2025, identified through PRISMA-guided screening, to examine how ChatGPT’s performance in evaluating students’ writing aligns or misaligns with established assessment principles. The selected studies were analyzed using a framework-based coding approach according to language assessment principles to synthesize patterns, trends, and gaps. Overall, the reviewed evidence suggests that ChatGPT tends to demonstrate greater consistency, particularly when assessing grammar and mechanics. However, concerns persist regarding fairness and contextual sensitivity in the evaluation of content and ideas, depending on assessment focus and context. Building on these findings, the review identifies several gaps in the existing literature, particularly the limited use of mixed-method designs that integrate perceptions, processes, and performance. Taken together, these strengths and limitations indicate that ChatGPT and human teachers play complementary roles in writing assessment. Accordingly, the proposed AI-human collaboration conceptual framework positions ChatGPT as a consistent surface-level evaluator and teachers as contextual judges, ensuring fairness, depth, and pedagogical value in AI-mediated writing assessment practices.