•  
  •  
 

Abstract

Artificial intelligence (AI)-supported assessment is increasingly used in education to automate scoring, provide feedback, predict learner performance, and support instructional decision-making. This systematic review synthesized evidence on the validity, fairness, and learning impact of AI-supported assessment across educational contexts. A systematic search of ERIC, Web of Science, Scopus, PsycINFO, PubMed, and Education Source identified peer-reviewed studies published between 2010 and 2025. Twelve studies met the inclusion criteria and were synthesized narratively following PRISMA 2020 guidance. The findings indicate that AI-supported assessment shows promising validity evidence, particularly in automated writing assessment, short-answer scoring, and predictive learning analytics, with several studies reporting moderate-to-strong alignment between AI-generated scores and human ratings or external performance indicators. However, validity evidence was often concentrated on score agreement, with limited attention to construct and consequential validity. Fairness findings were mixed; some systems showed limited additional bias, whereas others revealed subgroup disparities, unequal prediction, or compressed score distributions. Learning benefits appeared strongest when AI tools were used formatively to provide timely feedback, support revision, or identify at-risk learners. Overall, AI-supported assessment is promising but non-neutral, requiring rigorous validation, fairness monitoring, transparent reporting, and sustained human oversight.

Share

COinS