cs.LGSep 7, 2026

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

Authors: Jan CorazzaDaniil KaminskyiSimon LutzPatrick NossolHadi Partovi AriaZhe XuDaniel Neider

Abstract

We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.

Explore similar work

CardsList
  1. Reinforcement Learning with Temporal-Logic-Based Causal Diagrams

    Jun 23, 2023Yash Paliwal, Rajarshi Roy, Jean-Raphaël Gaglione +5Causal GraphAutomata