cs.AIOct 8, 2026

A 3D Characterization Framework for Intelligent Sequential Decision Making

Authors: Sadig Gojayev, Carolina Fortuna

Organizations: Department of Communication Systems, Jožef Stefan Institute, SI-1000 Ljubljana, Slovenia

Abstract

Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.

Figures & tables

Explore similar work

CardsList
  1. Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

    May 7, 2026Roger Creus Castanyer, Pablo Samuel Castro, Glen BersethAI Agent EvaluationAI Agent Benchmarks

  2. PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

    Sep 17, 2026Pyrros Koussios, Benjamin Jäger, John Hua Yao +3LLM EvaluationTemporal Reasoning in Language Models