Ultra Strong Machine Learning: LLM-Generated Explanations Do Not Yet Suffice for Teaching Humans Active Learning Strategy
Authors: Lun Ai, Johannes Langer, Ute Schmid, Stephen Muggleton
Organizations: Department of Computing, Imperial College London, UK · Faculty Information Systems and Applied Computer Science, University of Bamberg, Germany
Active learning is a general learning mechanism shared by artificial and human learners. Whether AI can teach humans such a strategy that transfers across domains is an open question. Ultra Strong Machine Learning (USML), a system whose explanations quantifiably improve human out-of-sample performance compared to self-learning, is uniquely positioned to answer this question. Prior USML work relied on hand-crafted explanation templates that require expert effort for each new domain and do not scale. We developed an explanation pipeline combining Inductive Logic Programming (ILP) with large language models (LLMs) to automate explanation generation and scoring. We tested whether these explanations achieve USML in a human trial teaching active learning strategies across three related domains. Our exploratory results show that concise, expert-written explanations benefit learners with higher initial performance, while pipeline-generated explanations provide no advantage over self-learning despite being rated as higher quality from an LLM-as-judge evaluation. This case study reveals a systematic gap that LLM quality metrics do not predict human learning outcomes. Our findings point to explanation complexity relative to task difficulty as a key factor, and call for explanation methods and evaluation criteria grounded in human cognitive constraints rather than LLM preference.
Figures & tables
Figure 1 : USML can quantifiably enhance human task performance compared to human self-learning from examples.
Figure 2 : Our explanation pipeline. ILP learns logic programs from examples; coding LLMs interpret programs and reasoning LLMs summarise their outputs into natural language explanations; LLM judges score candidates, optionally using expert-written references.
Figure 3 : The left block shows ILP-learned programs, where each episode is learned from a single circuit example. The middle block summarises relevant programs identified for the task. The right block is an action strategy based on relevant programs.
Figure 4 : Distribution of LLM judged scores for electric circuit domain explanations. RMs and CMs denote reasoning and coding LLMs, respectively. The significance of results has been highlighted by: p<0.05 (), p<0.01 (), p<0.001 ().
Figure 5 : Mean and standard error of human task performance by condition for high-baseline subgroups across domains.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : LLM-judged score distributions comparing pipeline-generated explanations against hand-crafted templates for game playing and algorithm discovery. Annotations and markers are consistent with Figure 4 in the main text.
Figure 7 : Visual domain introduction used in the study.
Figure 8 : Demonstration of AND gates used in the study.
Figure 9 : Worked-out example used for the introduction of the circuits domain in the study.
Figure 10 : Participants see this circuit without any highlighted nodes or edges during learning phase 1.
Figure 11 : Participants see these circuits without any highlighted nodes or edges during learning phase 2. These circuits contain the visual feedback for H2 . H1 sees no highlighted edges just as it is demonstrated in Figure 10 .
Figure 12 : Two traces for solving a simple linear circuit as seen by H2 . H1 does not see the group size annotations. Apart from that, the visual presentation is identical.
Figure 13 : Exemplary graphs used during the test phase.