Organizations: School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Guangdong, 518172, China. · Shenzhen Research Institute of Big Data (SRIBD), Shenzhen, 518172, China.
Many dynamical systems generate influences whose consequences are not fully exhausted in the realized trajectory at the moment they arise. Such consequences are often treated as absent, delayed or statically stored, leaving unclear how unrealized influence retains future relevance as the system evolves. Here we formulate causal-fate dynamics, in which generated influence may be realized, remain latent, or be transformed by subsequent dynamics, and give an exact finite-transport representation when the relevant maps are specified. A connectome-constrained Caenorhabditis elegans model first motivates the biological hypothesis that unresolved inter-neuronal influence may persist and contribute to later propagation; it does not establish such a mechanism in living animals. We next examine operational Internet routing, where a dynamically updated cross-observer history retains predictive information beyond the current local route state. We then use the representation to construct a Transformer architecture that explicitly transports and selectively realizes latent contextual influence while retaining language-modeling function. The three studies distinguish a model-motivated scientific hypothesis, an observational phenomenon compatible with future-relevant history and an executable construction for carrying unrealized influence through subsequent computation.
Figures & tables
Figure 1: Identical visible chess configurations can admit different futures. a, Persistent historical consequence. The visible piece placement and side to move are identical, but kingside castling is legal only when neither the king nor the rook has previously moved. Moving either piece and returning it to its original square does not restore the lost castling right. b, Accumulated historical consequence. The visible placement, side to move and move rights are identical, but a draw claim becomes available only when the same position has occurred for at least the third time. These rule-defined examples show that an identical visible configuration alone need not contain all history-dependent information relevant to future admissibility.
Figure 2: Selective realization of inter-neuronal influence in a connectome-constrained Caenorhabditis elegans model. From left to right, the fully realized reference immediately admits all generated cross-neuronal influence. In the selective-realization trajectory, intrinsic neuronal dynamics continue while unrealized influence is transported and transformed before selected components re-enter the realized trajectory. The local-autonomy counterfactual retains intrinsic neuronal evolution but excludes latent transport and subsequent cross-neuronal realization.
Figure 3: Different unfinished cross-observer histories distinguish futures from the same current local route state. Remote updates not yet matched by local observations form an ordered propagation front that evolves through addition, local matching and 300-s expiry, so identical local presents can retain different information about subsequent route evolution.
Figure 4: A Transformer architecture with finite-transport latent contextual influence and selective realization. a, Layerwise causal-fate construction. At each layer, every token maintains a realized state and a latent contextual state, whose sum forms the complete state. The complete Transformer map transports and transforms the latent contextual influence across depth, while a token-level realization rule determines when the accumulated influence enters the realized hidden state. b, Depth-dependent organization of realized contextual influence. The heat map shows the fraction of realized cross-token attention mass across Transformer layers and source–target context-distance bins on the held-out WikiText-2 test blocks, normalized within each layer over the distance bins.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Quantity
Value
Head neurons
188
Stimulus neurons with observations
168
Response neurons with observations
185
Observed off-diagonal pairs
23,264
Valid observed kernel pairs
23,180
High-confidence q<0.05 pairs
1,141
Appendix
Table S1: Fixed C. elegans data object. Here q denotes the multiple-testing-adjusted q value, and the empirical transform uses response amplitude a with fixed scaling constant s .
Metric
Selective
Local autonomy
Pair count
364
364
Relative propagation error
0.03515
1.0
Pearson correlation
0.999999991
not defined
Sign agreement
1.0
not defined
RMSE
0.007795
0.22176
Appendix
Table S2: Stimulus-stratified neural metrics.
Split
Samples
Positives
Complete episodes
Prefixes
Training
56,941
41,963
208
16
Validation
9,696
7,377
63
10
Held-out test
13,630
10,412
81
14
Appendix
Table S3: Chronological BGP analysis objects. Positives are local peer–prefix route changes within the subsequent 30 s.
Component
Definition and membership
Strong current state
Collector, remote observer and prefix identity; local peer AS; route presence, path length and identity bucket, origin, community count; latest local update type, age and change status; time-of-day sine and cosine. Shared by all extensions.
Local lags
Local update counts over 30 and 120 s, and announcement and withdrawal counts over 120 s.
Cumulative history
Remote update counts over 30 and 120 s, remote announcement and withdrawal counts over 120 s, and unmatched announcement–withdrawal balance.
Unordered history summary
Cumulative features plus unmatched count and distinct fingerprint count, current local–remote presence/path/origin mismatch, remote path length and identity bucket. Shared by all five-slot and content controls.
Five-event front
Type, age and path-identity bucket of the five newest unmatched events after consumption and expiry. Empty slots use type 0, age 3,600 s and path bucket 0.
Interpretable history summary
Unordered summary plus oldest/latest unmatched ages, announcement–withdrawal transitions within the queue, and latest ordinary remote-event type and age.
Appendix
Table S4: BGP representation components. All inputs are available at the sample timestamp.
Representation
Logistic loss
Gain
HGB loss
Gain
Current route without latest event
0.461229
-1.05%
0.383889
-13.83%
Strong current state
0.456428
0.00%
0.337255
0.00%
Ordinary local lags
0.460030
-0.79%
0.345884
-2.56%
Cumulative history summary
0.446902
2.09%
0.332171
1.51%
Interpretable history summary
0.455556
0.19%
0.315268
6.52%
Within-sample tuple shuffle (with ages)
0.457188
-0.17%
0.312173
7.44%
Appendix
Table S5: Held-out log loss and relative gain against the strong current-state model. All front slots are read from the eligible unmatched queue.
Model
Comparison
Gain
95% interval
Logistic
Five-event front vs current state
1.97%
[0.23%, 3.37%]
Logistic
Five-event front vs tuple shuffle
2.14%
[1.75%, 2.62%]
Logistic
Content order vs content shuffle
2.58%
[2.17%, 3.09%]
Logistic
Content order vs content inventory
0.96%
[0.41%, 1.67%]
Logistic
Five-event front vs unordered summary
0.53%
[-0.00%, 1.04%]
Logistic
Five-event front vs queue-removed summary
1.25%
[0.70%, 1.77%]
Appendix
Table S6: Paired continuous-time-block bootstrap comparisons (1,000 repetitions, ten 30-min blocks). Tuple shuffling retains event ages and is an encoding control; the content comparisons omit all per-event ages.
Figure S1: Corrected unmatched-front comparisons. Points show relative held-out log-loss gain and lines show paired 95% block-bootstrap intervals. The age-preserving tuple shuffle tests encoding; the two bottom comparisons omit per-event ages and preserve each sample’s event-content multiset.
Model
Time block
Samples
Positives
Current
Front
Gain
Logistic
1
312
244
0.4551
0.4541
0.21%
Logistic
2
1,225
1,034
0.3804
0.3805
-0.02%
Logistic
3
1,313
1,086
0.3800
0.3563
6.25%
Logistic
4
1,388
1,063
0.5029
0.5125
-1.90%
Logistic
5
1,600
1,258
0.4608
0.4434
3.78%
Logistic
6
1,764
1,332
0.4847
0.4671
3.63%
Appendix
Table S7: Five-event unmatched-front results by time block.
Model
Collector
Samples
Positives
Current
Front
Gain
Logistic
RV Chile
6,759
5,647
0.4192
0.3885
7.33%
Logistic
RIS rrc06
6,871
4,765
0.4930
0.5053
-2.50%
HGB
RV Chile
6,759
5,647
0.3231
0.2691
16.70%
HGB
RIS rrc06
6,871
4,765
0.3512
0.3404
3.06%
Appendix
Table S8: Five-event unmatched-front results by collector.
Model
Prefix
Samples
Positives
Current
Front
Gain
Logistic
164.163.52.0/22
346
20
0.6804
0.7174
-5.45%
Logistic
177.46.0.0/17
2,656
2,238
0.4155
0.5112
-23.02%
Logistic
186.83.117.0/24
1,078
631
0.6721
0.6746
-0.37%
Logistic
188.213.84.0/23
901
565
0.5778
0.5205
9.92%
Logistic
197.254.244.0/24
1,670
1,427
0.3243
0.2570
20.74%
Logistic
197.254.248.0/24
767
650
0.3406
0.2696
20.84%
Appendix
Table S9: Five-event unmatched-front results by prefix.
Seed
Model
Reference
Reference loss
Candidate loss
Gain [95% interval]
1729
Logistic
Current
0.456428
0.447414
1.97% [0.28%, 3.37%]
2718
Logistic
Current
0.456428
0.447414
1.97% [0.49%, 3.38%]
31415
Logistic
Current
0.456428
0.447414
1.97% [0.56%, 3.31%]
1729
HGB
Current
0.337255
0.305086
9.54% [7.02%, 11.37%]
2718
HGB
Current
0.342741
0.300615
12.29% [8.83%, 14.90%]
31415
HGB
Current
0.333338
0.302459
9.26% [5.49%, 12.15%]
Appendix
Table S10: Seed robustness of the five-event front against the current state. Intervals use 500 block-bootstrap repetitions.
Seed
Model
Reference
Reference loss
Candidate loss
Gain [95% interval]
1729
Logistic
Shuffle
0.459780
0.447934
2.58% [2.18%, 3.05%]
1729
Logistic
Inventory
0.452298
0.447934
0.96% [0.41%, 1.68%]
2718
Logistic
Shuffle
0.450398
0.447934
0.55% [0.32%, 0.83%]
2718
Logistic
Inventory
0.452298
0.447934
0.96% [0.40%, 1.72%]
31415
Logistic
Shuffle
0.455230
0.447934
1.60% [1.31%, 2.00%]
31415
Logistic
Inventory
0.452298
0.447934
0.96% [0.43%, 1.76%]
Appendix
Table S11: Seed robustness of age-free chronological content. Reference is shuffled content or the canonical inventory. Intervals use 500 block-bootstrap repetitions.
Model
Horizon (s)
Representation
Current
History
Gain [95% interval]
Logistic
30
Five-event front
0.456428
0.447414
1.97% [0.28%, 3.37%]
Logistic
30
History summary
0.456428
0.455556
0.19% [-1.60%, 1.61%]
Logistic
120
Five-event front
0.220203
0.206245
6.34% [2.56%, 11.10%]
Logistic
120
History summary
0.220203
0.208659
5.24% [1.60%, 9.42%]
HGB
30
Five-event front
0.337255
0.305086
9.54% [7.02%, 11.37%]
HGB
30
History summary
0.337255
0.315268
6.52% [3.96%, 8.77%]
Appendix
Table S12: Primary and secondary prediction horizons. Intervals use 500 repetitions.
Case
Type
UTC interval
Size pre
post
Net pre
post
Mismatch
Future
1
Add
00:00:01–00:00:04
1
2
1
2
1/1
1/1
2
Add
00:00:04–00:00:31
2
3
2
3
1/1
1/1
3
Add
00:00:31–00:00:35
3
4
3
4
1/1
1/1
4
Add
00:00:35–00:01:01
4
5
4
5
1/1
1/1
5
Add
00:01:01–00:01:04
5
6
5
6
1/1
1/1
6
Add
00:01:04–00:01:31
6
7
6
7
1/1
1/1
Appendix
Table S13: Illustrative unmatched-queue transitions; removal denotes consumption or expiry. Mismatch and future-change columns give before/after binary values.
Trajectory
Perplexity
Cross entropy
Token–layer realization ratio
Fully realized
38.6512
3.65458
1.0
Token-local
1315.24
7.18178
0.0
Selective
39.9138
3.68672
0.884022
Appendix
Table S14: Transformer test metrics.
Figure S2: Validation and test trade-off for the Transformer realization threshold.
Action-conditioned time-series forecasting requires accounting for how future actions and exogenous forcings influence multiple targets through partially observed effects with different delays and persistence. Direct conditioning leaves the evolution and target-specific influence of these effects implicit in the predictor, while static relational graphs specify connections without tracking evolving effects. This motivates representing future-driver influence through structured latent states that evolve over the forecast horizon and route information to individual targets. We introduce BeliefGraph-JEPA, a structured latent world model that factorizes driver influence into typed latent-effect states. These states are rolled forward under future drivers and routed through a graph to target-specific nodes, forming the predictive base of a joint-embedding predictive architecture. A capacity-controlled residual supplements this base with direct driver information. On four multi-target clinical, agricultural, environmental, and industrial systems, the framework outperforms a range of pretrained and supervised known-future-covariate baselines. Matched controls isolate latent dynamics, future rollout, graph routing, and residual capacity; future rollout and graph-first residual routing improve forecasting across all four systems.
Yue Li, Kangqi Ni, Zhen Tan +1
Carnegie Mellon University · University of North Carolina at Chapel Hill · Stevens Institute of Technology
Aggregate performance on continuous-time dynamic graphs (CTDGs) combines, in a single score, the portion attributable to known temporal regularities and the additional predictive power of neural models. This study separates the two at the query level. We construct a mechanism-constrained predictor that uses pair recurrence, recency and history position, renewal patterns, and short sequential transitions while learning the compatibility within each mechanism. Across four CTDG datasets, this predictor recovers a substantial portion of the performance of strong neural baselines, and the recovered performance quickly saturates with a small, dataset-specific set of explicit mechanisms. Neural residuals concentrate on queries for which the positive and negative candidates have similar mechanism-execution profiles. Allowing conditional interactions among mechanisms is more effective than simply reweighting their existing contributions. Conditioning the contribution of one mechanism on the execution state of another recovers 54.9-73.2% of the original neural-only queries and improves overall paired accuracy on all four datasets. Although the magnitude of the effect varies across datasets, these results show that the performance gap of neural CTDG models need not be treated solely as an opaque difference in representational capacity. At least part of the gap is localized to queries with similar candidate execution profiles and can be functionally explained by conditional coordination among known, low-dimensional mechanisms.
Minwoo Yu, Young-guk Ha
Smart Computing Laboratory, Department of Computer Science & Engineering, Konkuk University, Seoul 05029, Republic of Korea
Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
SiYuan Ma, Yiqin Luo, Zhangji +8
1Nanyang Technological University · 2Southern University of Science and Technology · 3Tianjin University +7