Automated Grading

Momentum

3 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 57

All topics
CardsList
  1. Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

    Oct 8, 2026Xing Zhang, Guanghui Wang, Yanwei Cui +4Automated EvaluationLLM Agent Self-Improvement

  2. Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

    Oct 7, 2026Zhifan Sun, Sebastian Gombert, Jannik Lossjew +6Automated GradingEducational Assessment

  3. Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring

    Oct 7, 2026Zhifan Sun, Sebastian Gombert, Fabian Zehner +3Automated GradingRubric-Based Evaluation

  4. Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

    Sep 24, 2026Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh +2Prompt SensitivityLLM-as-a-Judge

  5. ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals

    Sep 15, 2026Bowen Qin, Yi Xie, Yesheng Liu +1Reward HackingLanguage Model Generation Evaluation

  6. Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

    Aug 30, 2026Nishant Balepur, Paiheng Xu, Wei Ai +3LLM EvaluationMultiple-Choice Question Answering

  7. WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader

    Aug 6, 2026Boshui Chen, Huiping Liu, Shaolei ZhangRL for Code GenerationAutomated Software Testing

  8. AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

    Aug 1, 2026Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fröhlich +1Rubric-Based EvaluationAutomated Grading

  9. The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

    Jul 31, 2026Ilya MikhelsonEducational AssessmentMulti-Turn Dialogue Evaluation

  10. Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks

    Jul 21, 2026Lachlan McGinnessVLM EvaluationEducational Assessment

  11. From Pixel to Prognosis: Convolutional and GLCM Feature Fusion for Automated Four-Class Cataract Severity Classification

    Jul 20, 2026K. Mithra, Prem Kumar SanthanamMedical Image ClassificationAutomated Grading

  12. ICME 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing

    Jul 6, 2026Wei Sun, Weixia Zhang, Linhan Cao +30Industrial Anomaly DetectionIndustrial Inspection

  13. Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

    Jul 2, 2026Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard +2LLM-as-a-JudgeComputer Science Education

  14. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

    Jun 23, 2026Tian Zheng, Kai-Tai HsuLLM-as-a-JudgeHuman-in-the-Loop Evaluation

  15. Confidence-Aware Automated Assessment of Student-Drawn Scientific Models

    Jun 18, 2026Luyang Fang, Yingchuan Zhang, Jongchan Park +3Selective PredictionVision Transformer

  16. Evaluation of Image Matching for Art Skills Assessment

    Jun 18, 2026Asaad Alghamdi, Michael Poor, Trung-Nghia Le +1Image MatchingSiamese Neural Networks

  17. CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

    Jun 17, 2026Marco Becattini, Niccolò Caselli, Matteo Minin +2Software EngineeringMulti-Agent LLM Systems

  18. LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

    Jun 16, 2026Xiwei Xu, Chen Wang, Jacky Jiang +5LLM-as-a-JudgeEducational Assessment

  19. Semantic Grading of Written Answers in Low-Resource Language Bangla Using a Fine-Tuned Lightweight Language Model

    Jun 10, 2026Meherun Farzana, Aniket Joarder, Mahmudul Hasan +1Educational AssessmentLow-Resource Language Processing

  20. Towards Fully Automated Exam Grading: Fairness-Aware Recognition of Handwritten Answers with Foundation Models

    Jun 9, 2026Hartwig GrabowskiVLM EvaluationAlgorithmic Fairness

  21. RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

    Jun 8, 2026Yiteng Mao, Kenan Xu, Yijia Lyu +3Mathematical Reasoning BenchmarksLLM Evaluation

  22. Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations

    Jun 7, 2026Hartwig Grabowski, Michael CanzEducational AssessmentAutomated Grading

  23. Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

    Jun 7, 2026Qilin Zhou, Zhuo Wang, Yue Li +1LLM EvaluationLLM-as-a-Judge

  24. EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

    Jun 4, 2026Zhihao Wu, Linhai Zhang, Taiyi Wang +4LLM-as-a-JudgeRubric-Based Evaluation

  25. Deep Learning-assisted AMD Staging based on OCT and OCT Angiography

    Jun 3, 2026Yukun Guo, Tristan T. Hormel, An-Lun Wu +4Medical ImagingOptical Coherence Tomography