A 3D Characterization Framework for Intelligent Sequential Decision Making
Organizations: Department of Communication Systems, Jožef Stefan Institute, SI-1000 Ljubljana, Slovenia
Abstract
Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.
Figures & tables
| State space representation refers to how the mdp state space is specified in the sense of mdp . The levels of human prior are defined as follows: | |
| High | Humans explicitly design the state representation, fully specifying its structure and semantics, as in symbolic planners based on PDDL [ 27 ] . |
| Medium | Humans prescribe a high-level representational structure or template, while allowing the system to learn internal parameters or details, as in logical neural networks or partially structured encodings [ 17 ] . |
| Low | No explicit representational structure is imposed by humans; the model learns state representations end-to-end from interaction or data, as in fully end-to-end reinforcement learning systems trained via self-play [ 28 ] . |
| Mechanism refers to the algorithmic structure responsible for reasoning and inference. The levels of human prior are defined as follows: | |
| High | Humans design the inference procedure, such as logical rules or search heuristics. |
| Medium | Humans select the general type of mechanism to employ, while the algorithm handles the details; for example, in Value Iteration Networks [ 29 ] , humans prescribe a value-iteration–based planning structure, while the network learns the reward and transition representations end-to-end. |
| Knowledge Input concerns how training data, constraints, or reward structures are introduced, directly shaping the transition function and reward function in mdp formulation. | |
| High | Humans explicitly specify the training data, define relevant constraints, and design or shape reward structures, fully guiding how the model learns from examples, as in reinforcement learning with human-structured environments and hand-crafted training scenarios (e.g., Q-learning with expert-designed tasks). |
| Medium | Humans provide curated or loosely structured data without fully specified objectives, allowing the system to organize patterns and infer useful structure while still retaining partial human guidance, such as transformers trained on human-curated datasets. |
| Low | The model autonomously extracts training signals from large-scale interaction or self-generated experience, with minimal human specification beyond initial setup, as in self-supervised pretraining of large language models on unlabeled data [ 30 ] . |
| Learning reflects how the model updates its parameters in response to the training signal and the extent of human involvement in this process. | |
| High | Humans explicitly shape the optimization process through curated objectives, staged training procedures, or tightly controlled learning phases, as in reinforcement learning from human feedback (RLHF) for large language models. |
| Medium | Humans specify general optimization settings—such as learning rates, network architectures, or update rules—while the system autonomously manages the internal learning dynamics, as in Double Deep Q-Networks (DDQN) [ 31 ] . |
| Guidance captures the extent to which humans specify task details or constrain model behavior at runtime. | |
| High | Humans provide detailed task instructions, including explicit constraints, edge cases, and step-level guidance, often through fully specified prompts or rule descriptions, as in Least-to-Most prompting techniques [ 32 ] . |
| Medium | Humans specify high-level task constraints or goals without defining the complete execution structure, such as providing domain rules (e.g., “a smaller disk cannot be placed beneath a larger one” in toh ) while leaving reasoning to the model. |
| Low | Humans provide only minimal task information—such as specifying the start and goal states in graph-based solvers—while the system autonomously determines the reasoning process and action sequence. |
| Decision describes how actions or outputs are selected by the system during inference and the extent of human involvement in this process. | |
| High | Humans explicitly drive or supervise the decision-making process, often through interactive or step-wise instructions. The model depends on human-provided reasoning steps, clarifications, or corrective feedback to determine each action, as in interactive planning with large language models (e.g., DEPS) [ 33 ] . |
| Medium | Humans specify the decision framework or evaluation criteria—such as accuracy thresholds, selection rules, or preferred heuristics—while the system autonomously executes decisions within these constraints, as in safe reinforcement learning approach in constrained mdp [ 34 ] . |
| Model | State (S) | Action (A) | Transition (T) | Reward (R) | Goal (G) |
| datos | task prompt, code , generated feedback (token sequences) | Invoke coding or repair LLM once | Probabilistic token generation: | 1 = working code, 0 = non-working code | Verified succ and isgoal functions |
| autotos [ 23 ] | task prompt, function, verifier feedback (token prefixes) | Single LLM generation step | Probabilistic token generation: | 1 = working code, 0 = non-working code | Verified succ and isgoal functions |
| Neurosolver [ 18 ] | (symbolic disk–peg assignments) | Legal disk moves | Deterministic: , guided by goal-conditioned value signal | Goal reward; penalties for illegal or redundant moves | All disks stacked on target peg (optimal steps) |
| FBRL [ 16 ] | Bit string (disk positions) | Valid disk moves (bit changes) | Deterministic: | +1 valid, -1 invalid, +10 for goal | Bit string for disks on Peg C (binary pattern) |
| Method | Model | Architecture | Training | Inference | ||||
| Representation | Mechanism | Knowledge | Learning | Guidance | Decision | Validation | ||
| datos | DeepSeek + Mistral | High | High | |||||
| DeepSeek + Gemma | Medium | High | Medium | Medium | Medium | Low | Medium | |
| autotos [ 23 ] | DeepSeek | High | High | |||||
| Qwen | Medium | Medium | ||||||
| Yi | Medium | High | Medium | Medium | High | Low | Medium | |
| Model Setup | Disk | Att. / Steps | Tokens | Time (s) | Total E.Mem (MB) | Total O.Mem (MB) |
| datos | ||||||
| DeepSeek + Gemma | ||||||
| DeepSeek + Mistral | ||||||
| autotos [ 23 ] | ||||||
| DeepSeek | ||||||
| Yi | ||||||