Organizations: Tohoku University · Research and Development Center for Large Language Models, National Institute of Informatics · National Institute of Informatics · University of Tokyo · Osaka Kyoiku University
While recent studies have explored human behavior and preference simulation using large language models (LLMs), it remains unclear how well LLMs can simulate subjective evaluations from real learners in educational settings. We investigate this question using real learner evaluation data on feedback for high-school biology questions at both the group and individual levels. We compare performance with and without learner-specific information, such as personality traits and evaluation examples, across six models. Our results show that LLMs still have a limited ability to simulate learner evaluations. Providing learner profiles and examples improves score calibration and individual-level simulation, but more often fails to improve group-level consistency. These findings highlight the need to investigate which learner information and adaptation strategies are effective for learner preference simulation.
Figures & tables
Figure 1: Overview of the experiments and research questions. We first evaluate LLM simulation without learner-specific evaluation data (Exp. 1), and then examine performance using real learner evaluation data, such as profiles and evaluation examples (Exp. 2). At the group level, we investigate simulation performance (RQ1) the effects of Big Five traits (RQ2), and the contribution of learner evaluation data to group- and individual-level simulation (RQ3). Finally, we analyze feedback types that lead to disagreement between LLMs and learner evaluations (RQ4).
Feedback Prompt
Explanation
Normal
Provides indirect hints that encourage reasoning rather than directly presenting the correct answer. It is designed to be combined with the other five prompts.
Keywords
Emphasizes keywords that are necessary for reaching the correct answer.
Actionability
Provides concrete processes or steps needed to reach the correct answer.
Novelty
Introduces advanced content that goes beyond the scope of the target question, such as university-level knowledge.
Table 1: Overview of the six feedback prompts from the dataset of Furuhashi et al. (2026) , with each prompt’s name and explanation.
Spearman’s ρ↑
MAE ↓
Model
Base
+BF
Base
+BF
GPT-5
0.638
0.625
0.393
0.294
Gemini-3.5-flash
0.434
0.706
0.355
0.177
Gemini-2.5-flash
0.532
0.098
0.456
0.382
Gemma-4-31B-it
0.563
0.687
0.459
0.440
Qwen3-VL-32B-it
0.353
0.265
0.274
0.154
Table 2: Agreement between LLMs and learner group evaluations measured by Spearman’s ρ and MAE. Baseline achieves higher Spearman correlation in some cases, whereas adding Big Five information improves MAE.
Figure 2: Group-level simulation results across models for Spearman’s ρ ( 2(a) ) and MAE ( 2(b) ). Larger models often achieve strong Spearman’s ρ under the zero-shot Baseline condition. The effect of few-shot prompting on Spearman’s ρ varies by model scale. For MAE, Big Five trait conditions, especially +BF and +BF (Score), generally reduce errors relative to the Baseline. Few-shot prompting improves MAE for GPT-5, Gemini, and Gemma models, whereas it tends to increase errors for the Qwen3-VL models.
Figure 3: Individual-level simulation results across models for Spearman’s ρ ( 3(a) ) and MAE ( 3(b) ). For Spearman’s ρ , zero-shot settings often achieve correlations around 0.1, whereas Few-shot settings substantially improve performance across all cases. In many cases, conditions with learner profile information outperform the Baseline condition. For MAE, Few-shot settings consistently achieve lower errors than zero-shot settings across all models.
Figure 4: Confusion matrices between human and LLM evaluations (zero-shot Baseline setting). Each subplot compares rating distributions for correct and incorrect answers. Large-scale models tend to assign overly positive ratings for correct answers, whereas smaller models show stricter or middle-range ratings for incorrect answers.
Spearman ↑
MAE ↓
Model
All
Correct
Incorrect
All
Correct
Incorrect
GPT-5
0.712
0.567
0.651
0.427
0.468
0.392
Gemini-3.5-flash
0.516
0.328
0.498
0.359
0.406
0.337
Gemini-2.5-flash
0.563
0.445
0.468
0.473
0.489
0.462
Gemma-4-31B-it
0.533
0.566
0.445
0.482
0.501
0.468
Qwen3-VL-32B-it
0.271
0.375
0.289
0.347
0.530
0.240
Table 3: Spearman’s ρ and MAE for overall, correct-answer, and incorrect-answer groups under the zero-shot Baseline setting. Large-scale models show higher agreement for incorrect answers, whereas smaller models tend to show higher agreement for correct answers.
Total
Normal
Keywords
Actionability
Novelty
Coverage
Positivity
Model
Total
High
Low
High
Low
High
Low
High
Low
High
Low
High
Low
High
Low
GPT-5
65
65
0
7
0
3
0
11
0
6
0
29
0
9
0
Gemini-3.5-flash
68
61
7
8
0
2
0
10
0
4
2
26
0
11
5
Gemini-2.5-flash
94
89
5
13
1
4
0
18
0
8
2
32
0
14
2
Gemma-4-31B-it
84
83
1
12
0
2
1
15
0
7
0
32
0
15
0
Qwen3-VL-32B-it
55
54
1
6
0
3
1
7
0
4
0
23
0
11
0
Table 4: Distribution of large disagreement cases ( ∣ difference ∣≥1.0 ) across feedback types. Values indicate cases counts for each model. Because each feedback instance is evaluated on six criteria, disagreements are counted separately by criterion. High and Low indicate cases where the LLM assigns higher or lower ratings than humans.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Criteria
Total
Normal
Keywords
Actionability
Novelty
Coverage
Positivity
All Criteria
92
13
5
16
8
34
16
Guidance for Review
16
3
-
3
1
6
3
New Knowledge
22
2
2
4
2
9
3
Key Points Clarity
13
2
-
3
2
5
1
Ease of Understanding
18
2
1
3
1
7
4
Trustworthiness
8
2
1
-
1
3
1
Appendix
Table 5: Number of cases where group-level human evaluations differ from those of at least two LLMs by more than 1.0 points. In total, 92 instances are identified across all criteria. Disagreements occur most frequently for the Guidance for Review, Ease of Understanding, and New Knowledge criteria. LLMs also tend to assign higher ratings to information-rich feedback such as Coverage .
Feedback Type
Content
Criteria
Coverage
You selected the combination of heated enzyme solution and heated starch solution, but when enzymes are heated to 100°C, the protein denatures and the enzyme loses its catalytic activity. For digestion to proceed in this experiment, both the enzyme remaining active and the presence of starch as a substrate are required. For each of the four combinations a–d, check the activity state of the enzyme (whether heated or not) and the state of the starch (whether heated or not), and organize which conditions are essential for the digestion reaction. Taking into account the properties of enzymes (heat-induced denaturation), compare each combination once more. 加熱酵素液と加熱デンプン液の組合せを選んでいますが、酵素は100℃で加熱されるとタンパク質が変性して失活し、触媒機能を失います。この実験で消化が進むためには、酵素が活性を保っていることと基質であるデンプンが存在することの両方が必要です。四つの組合せa~dについて、酵素の活性状態(加熱の有無)とデンプンの状態(加熱の有無)をそれぞれ確認し、どちらの条件が消化反応に必須かを整理してください。酵素の性質(熱による失活)を踏まえて、もう一度各組合せを比較検討しましょう。
Key Points Clarity
Model
Evaluation Reason
Qwen3-VL-8B-it
The feedback identifies the core issue \CJK@punctchar \CJK@uniPunct 0"80"94 that enzyme denaturation and the state of starch both affect the reaction \CJK@punctchar \CJK@uniPunct 0"80"94 making the key point clear. However, it does not directly link the experimental results to the explanation. 酵素の失活とデンプンの状態が反応に影響するという核心を突いているのでポイントが明確になったが、実験の結果を直接結びつけて説明していなかった。
Qwen3-VL-32B-it
The key point that both enzyme activity and the presence of a substrate are required was clearly emphasized, making it easy to understand what needed to be learned. On the other hand, it would have been better if there had been a more detailed comparative explanation of why the other combinations were incorrect. 酵素の活性と基質の存在が両方必要であるというポイントが強調されていて、理解すべきことが明確になった。一方で、なぜ他の組合せが間違っているのかについての比較説明がもう少し詳しくあれば良かった。
Gemma-4-31B-it
The feedback directed me to organize the relationship between enzyme activity and the substrate, which clarified what I should focus on studying. However, it did not specify which pages of the textbook I should review. 酵素の活性と基質の関係を整理するように指示があったので、勉強すべき方向が見えました。一方で、教科書のどのページを読めばいいかまでは書いてありませんでした。
GPT-5
The core idea that both enzyme activity and the presence of a substrate are necessary for digestion was well emphasized, helping me organize what I needed to understand. However, the explanation lacked a direct conclusion about which combination \CJK@punctchar \CJK@uniPunct 0"80"94 a or b \CJK@punctchar \CJK@uniPunct 0"80"94 actually produces the most digestion. / The important criteria of "the enzyme remaining active" and "the substrate being present" were clearly laid out, making the basis for judgment explicit. However, there was no prioritization among options a–d as to which is optimal, making it difficult to feel confident in the final selection. 消化が進むためには酵素の活性と基質の存在が必要という核心が強調され、理解すべきポイントが整理できました。しかし、実際に最も進むのがaかbかなど、具体的な結論に直結する説明が不足していました。 / この問題で重視すべき『酵素が活性を保つこと』と『基質が存在すること』というポイントが整理され、判断基準が明確になりました。選択肢a〜dの中でどれが最適かの優先順位づけがなく、最終的な選択の確信が持ちにくいです。
Appendix
Table 6: Continued examples of feedback instances with large disagreements ((| difference | ≥ 1.0) between group-level learner evaluations and LLM evaluations. The table shows the feedback type, feedback examples, and the corresponding evaluation criteria. Notably, the Coverage feedback is characterized by comprehensively covering the steps leading to the correct answer; as a result, none of the models agreed with the human evaluations, as reflected in the evaluation reasons provided for each model.
Feedback Type
Content
Criteria
Coverage
When examining the combinations of causes and their corresponding results, you need to organize the mechanisms behind each environmental problem. This problem requires you to accurately understand the following four causal relationships: u (heavy metals and DDT) → c (biomagnification), e (chemical fertilizers) → e (eutrophication), o (factory emissions and exhaust gases) → b (acid rain), and ka (reduced groundwater infiltration) → f (depletion of spring water). In your selected combination, at least one of the causes u–ka has been incorrectly matched with its result, so please review your textbook to confirm the specific environmental impact of each cause (biomagnification, eutrophication, acid rain, and groundwater issues). In particular, try to reorganize how excessive use of chemical fertilizers leads to water quality problems, and how heavy metals accumulate through the food chain. あなたは各原因と結果の組合せを検討する際、それぞれの環境問題がどのようなメカニズムで発生するかを整理する必要があります。この問題では、ウ(重金属・DDT)→c(生物濃縮)、エ(化学肥料)→e(富栄養化)、オ(排煙・排ガス)→b(酸性雨)、カ(浸透水減少)→f(湧水枯渇)という4つの因果関係を正確に把握することが求められています。選択した組合せでは、原因ウ~カのうちどれかの結果が誤って対応づけられているため、各原因が引き起こす具体的な環境影響(生物濃縮、富栄養化、酸性雨、地下水問題)を教科書で再確認してください。特に化学肥料の過剰使用がもたらす水質問題と、重金属が食物連鎖を通じてどう蓄積されるかを整理し直してみましょう。
Trustworthiness
Model
Evaluation Reason
Qwen3-VL-8B-it
The information is accurate at the textbook level and highly trustworthy. However, no concrete suggestions for improvement are provided in response to the student’s incorrect answer. 教科書レベルの正確な情報に基づいており信頼性が高い。しかし、学生の誤答に対する具体的な改善策が提示されていない。
Qwen3-VL-32B-it
The explanation was consistent with the teacher’s and felt grounded in correct knowledge. However, it would have been more trustworthy if the reason why my chosen answer was wrong had been explained in a bit more detail. 先生の説明と一致しており、正しい知識に基づいていると感じた。ただ、自分の選択肢がなぜ間違っているのかの理由がもう少し詳しく説明されていればより信頼できた。
Gemma-4-31B-it
The explanation logically shows the path to the correct answer, making it convincing. However, it does not explain in detail why the other options are incorrect. 論理的に正解への道筋が示されており、納得感があります。とはいえ、他の選択肢がなぜダメなのかまでは詳しく書かれていません。
GPT-5
The causal relationships pointed out were textbook-appropriate and highly plausible, and I felt the content was reliable. However, since no sources or evidence were cited, I felt I needed to verify it myself just to be sure. / The correspondences indicated (u→c, e→e, o→b, ka→f) were consistent with the textbook and class content, and I found the explanation convincing. That said, since no sources were cited, I think someone encountering this material for the first time might feel the need to fact-check it. 指摘された因果関係は教科書的で妥当性が高く、内容は信用できると感じました。ただ、出典や根拠の提示がないため、念のため自分でも確認が必要だと思いました。 / 指摘された対応(ウ→c、エ→e、オ→b、カ→f)は教科書や授業内容と一致しており、内容に納得感がありました。ただ、出典や根拠の提示がないので、初めて学ぶ人には裏取りの必要があるようにも感じます。
Appendix
Table 7: Examples of feedback instances with large disagreements (| difference | ≥ 1.0) between group-level learner evaluations and LLM evaluations. This table shows the feedback type, feedback examples, and the corresponding evaluation criteria. Notably, the Model feedback tends to contain arrow ( → ) notation in its content. Furthermore, as reflected in the evaluation reasons, hallucinations were observed in some model outputs \CJK@punctchar \CJK@uniPunct 0"80"94 for instance, references to statements such as “the explanation was consistent with the teacher” despite no such information being present in the feedback.
Figure 5: Basic system prompt used as the baseline condition without learner profile information.
Figure 6: System prompt using Big Five personality trait information based on the questionnaire proposed by Murakami and Murakami (1997) , which consists of 70 items.
Figure 7: System prompt using score-based Big Five personality trait representations.
Figure 8: System prompt using pre-questionnaire information, including generative AI usage and basic biology-related items.
Figure 9: System prompt combining Big Five personality traits and pre-questionnaire information.
Figure 10: System prompt combining score-based Big Five personality traits and pre-questionnaire information.
Figure 11: Basic user prompt used as the Zero-shot baseline condition.
Figure 12: Few-shot user prompt using learners’ evaluations of other questions across six evaluation criteria.
Although large language models (LLMs) show promise for generating educational feedback, it remains unclear how feedback should be designed (e.g., tone and coverage) to support answer revision and learner evaluations across learner profiles. We define six feedback designs for multiple-choice biology questions, including a baseline design and five variants with additional feedback elements, and conduct an empirical study with 321 high school students. We evaluate feedback using immediate revision performance and six subjective evaluation criteria, and analyze differences in subjective evaluations across learner profiles based on personality traits. Our results show that presenting task-relevant information clearly is associated with better immediate revision performance and is favorably evaluated across learner profiles, while we observe descriptive differences in evaluation patterns, particularly for informational novelty and affective framing. These findings support further investigation of personalized LLM feedback design.
Momoka Furuhashi, Kouta Nakayama, Noboru Kawai +3
Tohoku University · Research and Development Center for Large Language Models, National Institute of Informatics · Osaka Kyoiku University +2
Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compares targeted feedback against two controls: misaligned feedback (targeting a different but plausible misconception) and generic feedback (only identifying answer is wrong). We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under contrastive controls. Across seven LLMs (4B-120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance. Further analyses reveal a sycophantic failure mode: models behave less like students with misconceptions but more like problem-solvers who treat any corrective signal as a cue to abandon the simulated belief and re-solve from internal knowledge. To address this, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward; SFT yields notable gains up to +0.56, and SFS-aligned RL provides more consistent improvements than preference optimization. Our results establish misconception faithfulness as a challenging yet trainable property, motivating a shift from static output matching toward interactive, belief-aware student modeling.
Heejin Do, Shashank Sonkar, Mrinmaya Sachan
ETH Zürich, ETH AI Center · University of Central Florida · ETH Zürich
In many applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. For example, in admissions, committees assess candidates on attributes such as test scores, GPA, and research experience to evaluate their overall fit for the program. Another example arises in medical care where clinicians use patient reports of symptoms to consider preliminary diagnoses and assess risks. Each setting involves mapping multiple criteria to an overall evaluation -- a process that reflects the evaluator's underlying preferences. We focus on the fundamental question of learning these preferences. Many applications of this problem make specific modeling assumptions on evaluator preferences that may be substantially violated in the real world. We make the minimal assumption that the preference function is coordinate-wise non-decreasing, which is reasonable in a large number of evaluation settings. We theoretically characterize the severity of model mismatch for many common assumptions and show that it can lead to significant issues for learning evaluator preferences and other important downstream tasks. We then present an algorithm for learning evaluators' preferences that is robust to model mismatch. We prove theoretically that our algorithm can learn any preference function without sacrificing performance when the linearity assumption holds. Evaluations of our algorithm with synthetic simulations and real-world data confirm its ability to learn preferences robustly and illustrate key aspects of LLM and human preferences.