4-Tensor Attention Model for Semantic Physical Reality
Organizations: IndigoWave, Center for Quantum Spacetime, Sogang University, 35 Baekbeom-ro, Mapo-gu, Seoul 04107, Republic of Korea
Abstract
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
Figures & tables
| Object | Shape | Index order |
|---|---|---|
| token_ids , | ||
| pad_mask | , true on a PAD key | |
| scores, attention weights | ||
| logits |
| Stage | Tensor | Storage shape |
|---|---|---|
| Step 1 | ||
| Steps 2–5, summed over heads | ||
| After both residual adds | ||
| Readout-boundary norm | ||
| Tied readout | logits | |
| Shifted targets |
| 4-tensor | Baseline | |
| Width of each head | full fiber, | |
| Head merge | sum, no output projection | concatenation, then |
| Attention parameters per layer | ||
| MLP parameters per layer | ||
| Token table and readout | , tied |
| Item | Value |
|---|---|
| Corpus | ROCStories, window , condition on four sentences |
| Split | train / validation |
| Vocabulary | ; baseline with [BOS] , [EOS] |
| Width | , ; baseline , |
| Dropout | on both residual branches |
| Optimizer | AdamW, learning rate , weight decay |
| seed | parameters | epoch | acc. (%) | ||||
|---|---|---|---|---|---|---|---|
| 2 | 2 | 43 | 122.2M / 175.9M | 4/9 / 30/39 | 5.562 / 5.873 | 16.8 / 7.3 | 0.311 |
| 4 | 2 | 43 | 172.5M / 175.9M | 6/11 / 45/50 | 5.554 / 5.701 | 16.0 / 10.5 | 0.147 |
| 4 | 3 | 42 | 256.4M / 226.2M | 6/11 / 29/35 | 5.531 / 5.667 | 15.4 / 11.1 | 0.136 |