cs.CLMar 16, 2026

Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability

Authors: Fan Huang, Haewoon Kwak, Jisun An

Organizations: Indiana University Bloomington / United States

Abstract

Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce moral reasoning trajectories, sequences of ethical framework invocations across intermediate reasoning steps, and analyze their dynamics across six models and three benchmarks. We find that moral reasoning involves systematic multi-framework deliberation: 55.4--57.7% of consecutive steps involve framework switches, and only 16.4--17.8% of trajectories remain framework-consistent. Unstable trajectories remain 1.29 times more susceptible to persuasive attacks (p=0.015). At the representation level, linear probes localize framework-specific encoding to model-specific layers (layer 63/81 for Llama-3.3-70B; layer 17/81 for Qwen2.5-72B), achieving 16.8--22.2% lower KL divergence than the step-prior baseline. Activation steering applied during generation moves the framework-consistency--accuracy relationship, widening it for Qwen2.5-72B and erasing it for Llama-3.3-70B, and a probe-space layer sweep bounds the attainable drift reduction at 6.7--8.9%. We further propose a Moral Representation Consistency (MRC) metric whose underlying framework attributions are validated by human annotators (mean cosine similarity = 0.859), and we report what an automated coherence rater does and does not establish about it.

Figures & tables

Appendix figures & tables44 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning

    Jun 13, 2026Ali Dasdan, Manan Shah, W. Russell Neuman +3Moral ReasoningTask Framing

  2. Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs

    Jun 10, 2026Elizaveta Tennant, Benjamin Henke, Anita Keshmirian +5Moral ReasoningLarge Language Model Reliability