AI Grading Tools in Schools Face Criticism Over Reliability
netzpolitik.org
Key Takeaways
- Research into AI-powered grading tools like FelloFish and Edaira indicates they are unsuitable for classroom use.
- These systems fail to provide consistent evaluations, often assigning significantly different grades to the same text across multiple checks.
- Adversarial testing revealed that the tools are easily manipulated and prone to favoring AI-generated text or nonsensical metatext over student work.
- Researchers warn that instead of reducing teacher workload, these tools create an additional burden by requiring manual verification to ensure fairness.
How the AI Tools Fail
- The software relies on Large Language Models (LLMs), which function stochastically rather than deterministically, leading to output variability that makes standardized grading impossible.
- FelloFish is intended for student feedback loops, while Edaira assists teachers with grading assignments. In trials, identical submissions saw grading swings of more than a full letter grade.
- Evaluation criteria often prioritize syntactical similarity to the model's own suggestions rather than evaluating the actual semantic content or quality of student improvements.
Policy and Academic Implications
- Researchers at the University of Osnabrück argue that relying on LLMs to handle grading fails to address the underlying structural workload issues facing educators.
- The expectation that AI increases objectivity is invalidated by its tendency to exhibit a "self-recognition bias," where AI-produced content is consistently rated more favorably.
- Critics suggest that funds currently earmarked for AI software licensing would be more effectively spent on hiring additional human staff to provide personalized, transparent student feedback.