Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges
Organizations: Department of Computer Science, University of Exeter, Harrison Building, Streatham Campus, North Park Road, Exeter, EX4 4QF, Devon, United Kingdom
Abstract
Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formulation, visual pretraining, supervision, decoder design, and evaluation code change simultaneously. The central contribution of this review is an evidence-aware design map that crosses task regime with the point at which language-derived information intervenes. We characterise task regimes along six axes. These axes organise the literature into five broad task families: single-action, sequence, object-interaction, cross-view, and planning-oriented settings. C1-C3 locate interventions in context construction, goal/intention modelling, and future decoding, while C4 is treated as an adjacent, emerging grounding/executability extension. Unlike a generic processing pipeline, the map links each intervention to an appropriate counterfactual, failure diagnosis, and permissible evidence claim. Supporting contributions include a protocol-level audit of Ego4D-LTA and EPIC-KITCHENS-100, a multidimensional evidence profile, and the Backbone-Aware Comparison and Ablation Protocol (BCAP). The unresolved EK-100 record is treated as a reporting-comparability case study and is not used as a leaderboard. Evidence for LLM benefits, goal ambiguity, and horizon effects is therefore formulated as testable hypotheses requiring matched validation, not as causal conclusions. The accompanying package contains the coded evidence, source locators, protocol metadata, and versioned catalogue used in the review.
Figures & tables
| Survey | Ver. | LLM focused | Anticip. focused | Design centric | Protocol aware | LLM vs. backbone |
|---|---|---|---|---|---|---|
| Tang et al. [ 4 ] | v1, 2025 | ✓ | no | no | no | |
| Zhong et al. [ 5 ] | v2, 04/2026 | ✓ | no | no | ||
| Kong & Fu [ 6 ] | 2022 | no | ✓ | no | no | |
| Lai et al. [ 7 ] | v1, 2024 | ✓ | ✓ | no | ||
| Hu et al. [ 8 ] | 2022 | no | no | no | ||
| Stergiou & Poppe [ 83 ] | 2025 | no | no |
| Symbol | Meaning |
|---|---|
| Observed video prefix | |
| Context-enriched representation | |
| Future action sequence (verb–noun pairs) | |
| Latent goal or high-level intention | |
| Grounding state (program/scene graph/trajectory) | |
| Feasible set of futures consistent with |
| Axis | Operational values | Examples | Required reporting | Interpretive consequence |
|---|---|---|---|---|
| Observation | Pre-action clip; recognised history; single frame; partial ongoing action; cross-view stream | EK-100; Ego4D-LTA; AAG; Ego4D-STA; DCPGN | Window length, FPS, observed labels, view and adaptation access | Fixes what evidence is available before any language component acts |
| Target | One label; action sequence; verb+noun+time/object; open-text step; executable plan | EK-100; Ego4D-LTA/STA; TrajPilot; Anticipate & Act | Target schema, sequence length, localisation or solvability requirement | Changes the output space and meaning of “correct” |
| Vocabulary | Closed; open; hybrid/retrieval-constrained | EK-100; Ego-Exo4D open vocabulary; prototype methods | Class inventory, unseen/rare-class policy, text-to-label mapping | Separates long-tail difficulty from genuine open-vocabulary generalisation |
| Horizon | Anticipation time ; discrete horizon ; real-time rollout | EK-100; Ego4D-LTA; streaming/planning studies | or , observation length , candidates , sampling policy | Prevents single-action and sequence protocols from being merged |
| Goal access | Provided; masked; inferred latent goal; semantic conditioning only; absent | Goal-conditioned planning; TrajPilot goal masking; AntGPT; PAR-VLA | Whether a goal is observed, inferred, or approximated by prototypes/text | Distinguishes latent-goal inference from semantic feature conditioning |
| Evaluation | MT5R; edit distance; mAP; calibration/coverage; solvability/success | EK-100; Ego4D-LTA/STA; Du et al.; planning boundary studies | Evaluator, split, version, subset and source locator | Metrics across task families are not interchangeable |
| Task family | C1 context | C2 goal/intention | C3 decoding | C4 adjacent grounding/planning |
|---|---|---|---|---|
| Single-action, closed vocabulary | Backbone, modalities, history | Optional semantic/intention conditioning | Class scoring | Usually not required |
| Long-term sequence forecasting | Observed-action/video history | Latent or prompted goal/subgoal | Autoregressive or structured sequence prediction | Feasibility/procedural constraints |
| Object-interaction anticipation | Actor/object context, localisation | Interaction intent or prototypes | Verb/noun/time/box prediction | Geometric or affordance consistency |
| Cross-view/test-time adaptation | View-normalised context and memory | Prototype/textual progression clues | Multi-label anticipation | Optional adaptation constraints |
| Planning/robotic boundary | Perceptual state and history | Explicit goal or subgoal | Plan/step generation | Solvability, dynamics, execution |
| Comp. | Intervention | Required counterfactual | Characteristic failure signature | Permissible evidence claim |
|---|---|---|---|---|
| C1 | Context or representation change | Hold the decoder and goal pathway fixed; match training data and evaluation code where possible | Context noise, missing evidence, modality mismatch, or weak temporal representation | Representation-level contribution under the stated matched controls |
| C2 | Goal, intention, or semantic-conditioning signal | Hold C1 and C3 fixed; remove, replace, or perturb the goal/intention signal | Goal hallucination, overconstraint, or collapse to an incorrect procedural branch | Goal-conditioning contribution, not a general LLM effect |
| C3 | Objective, dependency model, or decoder change | Hold context and goal inputs fixed; match candidate space and decoding budget | Repetition, exposure drift, weak action dependencies, or poor calibration | Decoding-stability or sequence-modelling contribution |
| C4 | Adjacent feasibility, grounding, or planning constraint | Use the same candidate generator without the check, constraint, or planner | Invalid, inconsistent, unsafe, or unreachable futures | Feasibility-sensitive contribution on an outcome that measures executability |
| Method | Core idea | C1 | C2 | C3 | C4 |
| Context construction (C1) | |||||
| TransFusion [ 55 ] | Language summaries of past context | — | — | ||
| PALM [ 29 ] | MMR-based context + prompt design | — | — | ||
| SAFT [ 30 ] | Sequential textual correction memory | — | — | ||
| M-CAT [ 31 ] | Dual-role text (past+future) | — | — | ||
| AAG [ 86 ] | Single-frame RGB/depth + action-history semantics | — | — | ||
| Method | Core idea | E2 | E3 | Relation to anticipation |
|---|---|---|---|---|
| SparseVLM [ 51 ] | Text-aware visual token pruning | — | Transfer mechanism only | |
| PruneVid [ 53 ] | Static/dynamic token disentanglement | — | Transfer mechanism only | |
| VTS [ 52 ] | Key-frame saliency+novelty pruning | — | Transfer mechanism only | |
| AdaCM 2 [ 54 ] | Goal-aware rolling memory compression | Transfer mechanism only | ||
| ReKV [ 50 ] | Sliding-window KV-cache retrieval | Transfer mechanism only | ||
| CLAM [ 78 ] | Cross linear attentive memory | Direct anticipation evaluation |
| Factor | Choices | Trade-off | Example |
|---|---|---|---|
| Modality | RGB vs. +flow/audio/HOI | Richer vs. heavier | M-CAT [ 31 ] , PALM [ 29 ] |
| Format | Labels / narrations / captions | Semantic vs. fast | TransFusion [ 55 ] |
| Scope | Window vs. full history | Rich vs. length limit | PALM ( ablation) |
| Retrieval | None / MMR / similarity | Diverse vs. costly | PALM (MMR ICL) |
| Spatial | Global / HOI hotspots | Precise vs. context loss | AFF-ttention! [ 44 ] |
| Text role | Past only / future target / both | Strong vs. cost | M-CAT [ 31 ] |
| Factor | Choices | Trade-off | Example |
|---|---|---|---|
| Goal representation | Latent state / textual intention / executable goal | Flexible vs. interpretable/executable | AntGPT, Ant.&Act |
| Inference pathway | Prompted LLM / VLM text LLM / explicit reasoning policy | Simplicity vs. visual grounding | AntGPT, ICVL, INSIGHT |
| Uncertainty handling | MAP goal / top- goals / posterior-conditioned decoding | Efficient vs. robust | AntGPT, Mascaró et al. |
| Coupling to decoder | Prompt-level conditioning / feature-level fusion / constrained decoding | Modular vs. tightly aligned | ICVL, AntGPT |
| Temporal granularity | Immediate next goal / multi-step subgoal / procedure-level objective | Local precision vs. long-horizon coherence | GP-AMS |
| Supervision source | Implicit anticipation loss / auxiliary intention supervision / structured task signal | Broad applicability vs. stronger bias | Mascaró et al., GP-AMS |
| Factor | Choices | Trade-off | Example |
|---|---|---|---|
| Output space | Closed labels / open text / hybrid verbalisation | Unambiguous vs. expressive | AntGPT, PALM |
| Decoding policy | Greedy / beam / temperature-based sampling | Stable vs. diverse | AntGPT, PlausiVL |
| Sequence granularity | Stepwise autoregressive / multi-token prediction | Flexible vs. globally coherent | VideoPlan |
| Plausibility control | None / penalty / compatibility or counterfactual constraints | Simple vs. constrained | PlausiVL |
| Auxiliary training signal | Anticipation only / +auxiliary tasks / +consistency losses | Cleaner objective vs. richer supervision | VideoPlan, PlausiVL |
| Label normalisation | Native closed-set output / post-hoc text-to-label mapping | Direct evaluation vs. mapping noise | ActionLLM, PALM |
| Factor | Choices | Trade-off | Example |
|---|---|---|---|
| Grounding state | Symbolic program / scene graph / trajectory or latent dynamics | Interpretable vs. expressive | LEAP, SymAnt, HWM |
| Constraint source | Hand-crafted / learned / hybrid | Reliable vs. scalable | SymAnt, TR-LLM |
| Constraint timing | Post-hoc filtering / decoder-time conditioning / joint planning | Modular vs. tightly coupled | PlausiVL, LEAP, Ant.&Act |
| Planner type | None / symbolic planner / hierarchical MPC | Simplicity vs. stronger executability | Ant.&Act, HWM |
| Evaluation target | Standard anticipation metrics / logic or solvability checks / task success | Comparable vs. realistic | LEAP, Ant.&Act, HWM |
| Transfer burden | Domain-specific schemas / reusable priors / embodied adaptation | Strong control vs. portability | SymAnt, TR-LLM |
| Method | Ver. | Verb ED | Noun ED | Action ED | Primary-source location | |||
|---|---|---|---|---|---|---|---|---|
| Mascaró et al. [ 35 ] | v1 | 6 | 20 | 5 | 0.741 | 0.739 | 0.930 | Table 1; test set |
| AntGPT [ 22 ] | v1 | 8 | 20 | 5 | 0.6584 | 0.6546 | 0.8814 | Table 6; test set |
| PALM [ 29 ] | v1 | 8 | 20 | 5 | 0.6559 | 0.6401 | 0.8613 | Table 1; test set |
| AntGPT [ 22 ] | v2 | 8 | 20 | 5 | 0.6503 | 0.6498 | 0.8770 | Table 6; test set |
| PALM [ 29 ] | v2 | 8 | 20 | 5 | 0.6471 | 0.6117 | 0.8503 | Table 1; test set |
| Method | Year | Model / encoder config. | Family | Action MT5R | Publication status | Comparability |
| InAViT [ 26 ] | 2024 | 160M | Non-LLM, supervised | 25.8 | Archival | Level (i) |
| Video-LLaMA [ 27 ] | 2023 | 7B | LLM-augmented | 26.0 | Archival | Level (i) |
| PlausiVL [ 23 ] | 2024 | 8B (Q-Former + LLM) | LLM-augmented | 27.6 | Archival | Level (i) |
| V-JEPA 2 [ 24 ] | 2025 | ViT-g384 (1B, frozen) | Non-LLM, dense SSL | 39.7 | Preprint † | Level (i) |
| V-JEPA 2.1 [ 25 ] | 2026 | ViT-G (2B, frozen) | Non-LLM, dense SSL | 40.8 | Preprint † | Level (i) |
| † Preprint at the time of writing; interpret as emerging rather than settled evidence. | ||||||
| Benchmark | Metric | Horizon | Label space | Subset | Modality |
|---|---|---|---|---|---|
| Ego4D LTA v1 | ED / seq-F1 | ✓ | v1 | — | ✓ |
| Ego4D LTA v2 | ED / seq-F1 | ✓ | v2 | — | ✓ |
| EK-100 (official) | action MT5R | fixed | ✓ | ✓ | ✓ |
| EK-100 (@1s) | Action@1s | fixed | ✓ | ✓ | ✓ |
| EK-55 | mAP | fixed | ✓ | ✓ | ✓ |
| Assembly101 | task-specific | ✓ | ✓ | — | ✓ |
| Setting | Horizon | Ambiguity proxy | Constraint | Backbone | Evidence | Evidence profile | Source |
|---|---|---|---|---|---|---|---|
| LTA, goal-ambiguous | High | Low | Moderate | AntGPT (Ego4D-v1 val): inferred-goal conditioning reduces verb/noun ED from 0.735/0.753 to 0.724/0.744; increasing from 1 to 8 reduces ED from 0.734/0.748 to 0.707/0.719 | Archival; dedicated; direct | [ 22 ] | |
| Procedural + constraints | Any | Med. | High | Any | LEAP: action-anticipation accuracy on EPIC-KITCHENS 14.64% 16.98% (author ablation, not MT5R) | Preprint; partial; metric-mismatched | [ 39 ] |
| Rare/long-tail verbs | Any | High | Low | Any | M-CAT (EK-100 val): adding action/object text raises action Recall@5 from 18.4 to 23.7 overall and from 16.0 to 21.0 on tail classes (author ablation) | Preprint; dedicated; direct | [ 31 ] |
| Short-horizon, scaled backbone | s | Low | Low | High | Unresolved reporting incompatibility in this Level-(i) matrix: PlausiVL (8B) 27.6, V-JEPA 2 39.7, and V-JEPA 2.1 40.8 action MT5R on EK-100 val | Mixed maturity; cross-paper; Level (i) | [ 24 , 25 , 23 ] |
| Streaming, tight latency | Any | Any | Low | Any | STREAMMIND reports up to 100 FPS in its stated A100 setup; input resolution, batching, output regime, and end-to-end anticipation equivalence are not established here | Preprint; no anticipation isolation; transfer | [ 47 ] |
| Problem regime | Primary component(s) |
|---|---|
| Short-horizon, closed vocabulary ( s) | A scaled C1 backbone should be evaluated as a required diagnostic baseline before attributing gains to language augmentation |
| Long-horizon, goal-ambiguous ( ) | Evaluate C2 + C3; goal ambiguity is currently a heuristic proxy |
| Safety- or executability-critical | Evaluate C4 grounding; evidence remains emerging and domain-specific |
| Real-time / edge deployment | E2 + E3: token pruning with event-gated streaming |
| Long-tail and compositional generalisation | C1 + C2: open-vocabulary context with structured goals |
| Calibration-critical applications | C3 with post-hoc calibration (BCAP Step 6) |
| Question | Evidence-supported answer | Confidence | Principal remaining gap |
|---|---|---|---|
| Q1: What components define the design space? | C1–C3 form the core analytical pipeline; C4 is a useful but less mature grounding/executability extension. The decomposition is non-unique and methods may span components. | Moderate–high | Primary assignment is non-unique and was not blindly recoded in full by an independent second coder; row-level coding and the alternate-assignment sensitivity analysis are supplied for audit. |
| Q2: Which design factors matter? | Context fidelity, goal conditioning, decoding stability, and feasibility constraints recur across methods. Their relative value depends on horizon, task structure, and deployment budget. | Moderate | Few studies isolate one factor while holding backbone, data, and compute fixed. |
| Q3: When do language components add value beyond a backbone? | Author-controlled long-horizon studies support some language-related components. For EK-100, published tables contain a large unresolved discrepancy under nominally similar labels, so the present record cannot support a cross-family performance ordering. | Low–moderate | No identified study performs a pretraining-matched, code-verified comparison across horizons. |
| Claim | Principal studies | Groups | Shared evaluator | Controlled evidence | Confidence |
|---|---|---|---|---|---|
| Goal/intention conditioning can improve long-horizon anticipation | AntGPT, ICVL, INSIGHT, GP-AMS | 4 | No across papers | Dedicated or partial author ablations | Moderate |
| Semantic/action-history context can replace part of dense video in selected procedural settings | TransFusion, AAG/AAG+, PALM | 3 | No across papers | Modality/history ablations | Moderate, task-specific |
| A geometric intent signal can outperform textual conditioning in the studied regime | TrajPilot | 1 | Within study | Direct conditioning contrast | Moderate within study; unreplicated |
| Anti-repetition or structured decoding can improve sequence stability | PlausiVL, VideoPlan, AGA | 3 | No across papers | Heterogeneous author ablations | Low–moderate |
| Published EK-100 numbers are difficult to reconcile | PlausiVL, V-JEPA 2/2.1 and transcribed baselines | Multiple | No | No matched cross-family ablation | High for incompatibility; low for ranking |
| Grounding/verification may improve feasibility | LEAP, FactCheck, SymAnt and boundary cases | Multiple | No | Mainly within-study or non-benchmark evidence | Low/emerging |
| BCAP item | AntGPT [ 22 ] | PlausiVL [ 23 ] | V-JEPA 2.1 [ 25 ] |
|---|---|---|---|
| Evaluation implementation/version | Partial: an official repository and LTA inference path were identified; commit and file locators are recorded in Supplementary Table S7B, but the survey row was not rerun. | No versioned public implementation was identified in this audit; the status is therefore based on the paper description only. | Partial: an official repository, EK-100 evaluation script, and model configuration were identified and frozen in Supplementary Table S7B; the reported row was not rerun. |
| Pretraining disclosure | Partial: constituent checkpoints are named; exposure is not harmonised with comparison systems. | Partial: backbone/checkpoint information is reported; data volume is not matched to the non-LLM encoders. | Substantial disclosure of model/data scaling, but it remains unmatched to the language-augmented systems. |
| Prompt or decoding reproducibility | Partial: prompting strategy and variants are described; a complete version-frozen run configuration is not available for every row. | Partial: plausibility and decoding mechanisms are described; complete cross-paper run settings are not standardised. | Not applicable to an LLM prompt; attentive-probe and evaluation configuration still require versioned reporting. |
| Repeated stochastic evaluation | No uncertainty interval from repeated LLM evaluations identified. | No BCAP-style repeated-run uncertainty interval identified. | No repeated-run interval attached to the reported EK-100 survey row. |
| Deployment cost | Not reported in the full BCAP hardware/memory/token-cost format. | Not reported in the full BCAP hardware/memory/token-cost format. | Compute is discussed for the model family, but not as a matched per-example BCAP deployment row. |
| Permitted inference | Descriptive system comparison; author ablations support selected within-system component claims. | Descriptive system comparison; author ablations support plausibility and repetition-control claims. | Descriptive evidence for a densely pretrained C1 system; no causal claim about removing an LLM from a matched architecture. |