Ultra Strong Machine Learning: LLM-Generated Explanations Do Not Yet Suffice for Teaching Humans Active Learning Strategy
Authors: Lun Ai, Johannes Langer, Ute Schmid, Stephen Muggleton
Organizations: Department of Computing, Imperial College London, UK · Faculty Information Systems and Applied Computer Science, University of Bamberg, Germany
Active learning is a general learning mechanism shared by artificial and human learners. Whether AI can teach humans such a strategy that transfers across domains is an open question. Ultra Strong Machine Learning (USML), a system whose explanations quantifiably improve human out-of-sample performance compared to self-learning, is uniquely positioned to answer this question. Prior USML work relied on hand-crafted explanation templates that require expert effort for each new domain and do not scale. We developed an explanation pipeline combining Inductive Logic Programming (ILP) with large language models (LLMs) to automate explanation generation and scoring. We tested whether these explanations achieve USML in a human trial teaching active learning strategies across three related domains. Our exploratory results show that concise, expert-written explanations benefit learners with higher initial performance, while pipeline-generated explanations provide no advantage over self-learning despite being rated as higher quality from an LLM-as-judge evaluation. This case study reveals a systematic gap that LLM quality metrics do not predict human learning outcomes. Our findings point to explanation complexity relative to task difficulty as a key factor, and call for explanation methods and evaluation criteria grounded in human cognitive constraints rather than LLM preference.
Figures & tables
Figure 1 : USML can quantifiably enhance human task performance compared to human self-learning from examples.
Figure 2 : Our explanation pipeline. ILP learns logic programs from examples; coding LLMs interpret programs and reasoning LLMs summarise their outputs into natural language explanations; LLM judges score candidates, optionally using expert-written references.
Figure 3 : The left block shows ILP-learned programs, where each episode is learned from a single circuit example. The middle block summarises relevant programs identified for the task. The right block is an action strategy based on relevant programs.
Figure 4 : Distribution of LLM judged scores for electric circuit domain explanations. RMs and CMs denote reasoning and coding LLMs, respectively. The significance of results has been highlighted by: p<0.05 (), p<0.01 (), p<0.001 ().
Figure 5 : Mean and standard error of human task performance by condition for high-baseline subgroups across domains.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : LLM-judged score distributions comparing pipeline-generated explanations against hand-crafted templates for game playing and algorithm discovery. Annotations and markers are consistent with Figure 4 in the main text.
Figure 7 : Visual domain introduction used in the study.
Figure 8 : Demonstration of AND gates used in the study.
Figure 9 : Worked-out example used for the introduction of the circuits domain in the study.
Figure 10 : Participants see this circuit without any highlighted nodes or edges during learning phase 1.
Figure 11 : Participants see these circuits without any highlighted nodes or edges during learning phase 2. These circuits contain the visual feedback for H2 . H1 sees no highlighted edges just as it is demonstrated in Figure 10 .
Figure 12 : Two traces for solving a simple linear circuit as seen by H2 . H1 does not see the group size annotations. Apart from that, the visual presentation is identical.
Figure 13 : Exemplary graphs used during the test phase.
Deep active learning has previously been explored for LLM in-context sample selection, but not with methods that utilise recent advances in understanding of transformer activations. In this paper, we test the hypothesis that model activations could provide a fine-grained signal to optimise the selection of in-context examples. We present a comprehensive analysis of MLP activation-based deep active learning methods applied to in-context learning, including how different attention masking strategies impact active learning across diverse classification and generative datasets, using both Llama-3.2-3B and Qwen2.5-3B base models. However, we find a negative result: MLP and embedding layer outputs, viewed through the lenses of massive activations or the first four moments, do not correlate with example quality or task performance. Specifically, the absolute Spearman correlation coefficient is at most 0.33 for all tasks and models we tested, showing that such activation-based sampling should not be used for in-context learning. We hypothesise that this may be due to superposition, whereby models represent more features than they have dimensionality, suggesting that methods like Sparse Autoencoders (SAEs) may be a promising future direction.
Yaseen M. Osman, Geoff V. Merrett, Stuart E. Middleton
School of Electronics and Computer Science, University of Southampton, United Kingdom
Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations. Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior. However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an open question. In this opinion paper, we argue that self-explanations can be highly plausible, questionably faithful, and yet highly actionable. From a traditional XAI perspective, we identify the limitations of standard evaluation protocols for LLM-generated self-explanations and propose practical guidelines for assessing their plausibility and faithfulness. Moreover, we argue that evaluation should extend beyond these criteria to actionability, highlighting applications of LLM rationalization capabilities that support informed decision-making and appropriate action across diverse stakeholders.
When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise. In this context, explanations that respond to users' questions are crucial to improve their understanding of potential solutions and increase their trust in the system. To enable natural interaction with such a system, we present a multi-agent Large Language Model (LLM) architecture that is agnostic to the explanation framework and enables user- and context-dependent interactive explanations. We also describe an instantiation of this framework for goal-conflict explanations, which we use to conduct a user study comparing the LLM-powered interaction with a baseline template-based explanation interface.
Guilhem Fouilhé, Rebecca Eifler, Antonin Poché +2
IRIT, Toulouse, France · LAAS-CNRS, Toulouse, France · IRT Saint Exupery, Toulouse, France +2