Organizations: National University of Defense Technology · Nanyang Technological University · The Hong Kong Polytechnic University · Southeast University
Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional context and inference costs and makes the acquired learner-specific simulation capability difficult to reuse across different LLMs. We therefore propose Learner2Skill, which externalizes the simulation capability acquired from historical interactions into a persistent and reusable Simulation Skill. The Skill captures the learner's current learning state and recurring response patterns, evolves as new real interactions arrive, and can be adapted to a new LLM through lightweight executor calibration without reconstructing the learner from scratch. Experiments show that Learner2Skill more faithfully reproduces fine-grained learner behavior while reducing overall token cost, and that the same constructed Skills can be effectively reused across different LLM executors.
Figures & tables
Figure 1 : Comparison of (a) existing LLM simulation agents that progressively model learners from interaction histories, and (b) Learner2Skill, which constructs reusable Simulation Skills for cross-executor simulation and online evolution.
Figure 2 : Overview of Learner2Skill. (a) Construct learner-specific Skills from real histories. (b) Calibrate a shared execution protocol through cross-learner diagnosis and held-out validation. (c) Update learner states and response patterns from new real interactions, with model parameters and the calibrated protocol fixed.
Simulation Effectiveness ( ↑ )
Simulation Cost ( ↓ )
Cost-Effectiveness ( ↑ )
Method
Attempt
Process
Final Ans.
Realized Corr.
Outcome Acc.
Consistency
Acq.
Test
Total
CE
Supervised Learner-Performance Prediction
KES
–
–
–
–
52.92±0.10
–
–
–
–
–
DKVMN
–
–
–
–
61.16±0.10
–
–
–
–
–
SAKT
–
–
–
–
61.01±0.09
–
–
–
–
–
DAISim
–
–
–
–
62.63±0.05
–
–
–
–
–
Table 1 : Main results on LearnerTrace-2K. Effectiveness results are reported as mean ± standard deviation over three runs. All LLM-based methods use Claude Sonnet 4.6 as the simulation executor. Higher is better for simulation effectiveness and cost-effectiveness, while lower is better for simulation cost. “–” indicates that the metric is not applicable. The best result is shown in bold and the second-best result is underlined .
Configuration
Simulation Fidelity ( ↑ )
Simulation Cost ( ↓ )
Setting
Executor
Attempt
Process
Final Ans.
Realized Corr.
Outcome Acc.
Consistency
Build
Calib.
Test
Total
Initial
Claude Sonnet 4.6
99.64
54.19
40.13
57.71
63.43
71.54
125.36M
57.23M
116.57M
299.16M
Reuse
Gemini 3.8 Flash
99.78
54.53
43.74
57.22
63.73
75.21
0
60.55M
132.13M
192.68M
Reuse
Qwen3.8 Flash
98.27
50.15
40.07
55.24
59.77
72.21
0
53.21M
103.12M
156.33M
Reuse
DeepSeek V4 Flash
99.94
48.93
43.11
57.42
60.86
75.22
0
60.17M
80.60M
140.77M
Table 2 : Cross-executor reuse of Simulation Skills on LearnerTrace-2K. Initial denotes the complete pipeline used to construct the Skills, whereas Reuse directly applies the constructed Skills to a new executor without repeating Skill Construction.
Method
Attempt ↑
Process ↑
Final Ans. ↑
Realized Corr. ↑
Outcome Acc. ↑
Consistency ↑
Full Learner2Skill
99.64
54.19
40.13
57.71
63.43
71.54
w/o Learning State
99.13 -0.51
56.80 +2.61
40.03 -0.10
53.80 -3.91
62.03 -1.40
70.51 -1.03
w/o Response Patterns
99.98 +0.34
35.05 -19.14
39.64 -0.49
58.50 +0.79
62.60 -0.83
70.18 -1.36
w/o Executor Calibration
99.63 -0.01
53.46 -0.73
39.75 -0.38
57.21 -0.50
62.57 -0.86
71.40 -0.14
w/o Skill Evolution
99.53 -0.11
52.83 -1.36
39.38 -0.75
56.86 -0.85
62.51 -0.92
71.37 -0.17
Table 3 : Ablation study of Learner2Skill on LearnerTrace-2K. Colored numbers indicate changes relative to the full model.
Figure 3 : CAT performance with generated data. Numbers above the lines denote F1 improvements after augmentation.
Agent skills are commonly distributed as SKILL.md files: human-readable procedural documents that describe workflows, tools, resources, and domain conventions. While convenient for inspection and reuse, this design requires the same reusable procedure to be repeatedly injected into the runtime context. We propose Skill-to-LoRA(S2L), a behavior-centric skill representation that replaces runtime skill text with skill-specific LoRA adapters. Rather than compressing the skill document itself, S2L models the behavioral change induced by the skill text: offline, the complete SKILL.md is used to synthesize skill-guided demonstrations; online, the full document is omitted and the corresponding LoRA adapter is dynamically loaded to activate the learned skill behavior. We evaluate S2L with Qwen3.6-27B on a 21-skill subset of SWE-Skills-Bench. Compared with the no-skill and Full Skill Text baselines, S2L improves pass rate by 2.9 and 5.2 percentage points, respectively, while reducing per-step token cost by 6.6% relative to Full Skill Text prompting. S2L matches or improves Full Skill Text on 18/21 skills and the no-skill baseline on 15/21 skills. Control experiments further show that the gains depend on skill-specific adapter alignment: Wrong-LoRA and Shared-LoRA both reduce performance. These results suggest that many procedural agent skills can be converted from runtime instructions into trainable, dynamically loadable behavioral modules. Code will be released upon acceptance.
Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.
Rose Niousha, Minwoo Kang, Narges Norouzi
Department of Electrical Engineering and Computer Sciences University of California, Berkeley
Skills have become the de facto way to enable LLM agents to perform complex real-world tasks with customized instructions, workflows, and tools, but how to learn them automatically and effectively remains unclear. We introduce SkillLearnBench, the first benchmark for evaluating continual skill learning methods, comprising 20 verified, skill-dependent tasks across 15 sub-domains derived from a real-world skill taxonomy , evaluated at three levels: skill quality, execution trajectory, and task outcome. Using this benchmark, we evaluate recent continual learning techniques, those leveraging one-shot, self/teacher feedback, and skill creator to generate skills from agent experiences. We find that all continual learning methods improve over the no-skill baseline, yet consistent gains remain elusive: no method leads across all tasks and LLMs, and scaling to stronger LLMs does not reliably help. Continual learning improves tasks with clear, reusable workflows but struggles on open-ended tasks, and using stronger LLM backbones does not consistently produce better skills. Our analysis also revealed that multiple iterations in continual learning facilitate genuine improvement via external feedback, whereas self-feedback alone induces recursive drift. Our data and code are open-source at https://github.com/cxcscmu/SkillLearnBench to enable further studies of automatic skill generation and continual learning techniques.