Contextual trajectory and incremental contextual displacement: Towards using LLMs to understand dynamic, utterance-specific meaning construction
Authors: Grayson Wycliffe Storer, Julia Witte Zimmerman
Organizations: Department of Mathematics and Statistics. · Department of Computer Science, University of Vermont, Burlington, 05405, Vermont, United States. · Computational Story Lab. · Computational Ethics Lab, Vermont Complex Systems Institute, Burlington, 05405, Vermont, United States.
Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context. We construct token-wise incremental trajectories by repeatedly recomputing a token's CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds. We evaluate this approach using garden-path sentences as a test case with characteristic features. Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls. We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type. We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales. In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena. Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.
Figures & tables
Figure 1: An illustration of characteristic semantic gap between garden path and equivalent disambiguated sentences.
Figure 2: An illustration of incremental context, as opposed to layerwise LLM analysis.
Figure 3: Cosine similarity and Euclidean distance between garden path and disambiguated sentences, all having a critical region occurring at the same point, specifically the 3rd word in the sentence. We see the same behavior that we describe in our formalization: the two trajectories are identical at the beginning of the sentence, the largest separation between the trajectories occurs in the critical region (2nd index, 3rd word), and the semantic similarity increases through the last increment.
Figure 4: Three different jokes of varying complexity exhibiting similar global “U-turn” behavior. From left to right, their transcripts are as follows: “What do you get if you cross a dog with Penn & Teller? Two Labracadabradors.”, “What do you call cheese made with Nirvana music? Curd Cobain!”, “Did you hear they’re making a vehicle that has a built-in gym? It’s a muscle car.”
Figure 5: Flow charts depicting emergent organization in latent space. The trajectories in the upper 4 plots are colored according to the sentence type and the trajectories in the lower 4 plots are colored according to the sense of the polysemous token. Note that this CLS token is not trained separately, but incorporates information from the other tokens in the context.
Appendix figures & tables35 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: These charts are not particularly meaningful in and of themselves, but we included them as an example of the caveats we are bringing up. First, architecture matters to epistemology: In the RoBERTa model, the attention process compresses position along with all other linguistic information accessible to the model. In DeBERTa, position is stored separately, leading to different behaviour in the representations between the two models as (architectural) context is added, even though both models do share significant aspects of their behaviour. Second, because (broadly-construed) architecture matters, we cannot assume our experience of language (via any representation: formal, informal, etc.) will map cleanly to what happens in LLMs, even though the language we both generate has significant commonalities. As what is essentially random gibberish is added to the (architectural) context, we doubt most people would perceive much difference between “… (234 tokens of gibberish) … bank” and “… (240 tokens of gibberish) … bank”; however, the model continues to adjust its representations. These two graphs show when the bank token is anywhere in the context (left) versus when it is at the end (right). The “context” here is a toy context: just random other tokens from the model’s vocabulary. Additionally, as we were hoping would be the case, DeBERTa, the model with disentangled attention does not obviously show a correlation between context length and magnitude of vector change, compared to RoBERTa.
Figure 7: Every token in a sentence has its own respective trajectory for the duration that it is in the sentence. This plot shows the PCA decomposition of the trajectory of every single token in a garden path sentence, plotted in the same PCA space. In this case the sentence is “The committee mentioned the issue would cause a problem at the meeting.”
Figure 8: PCA plots of the trajectory of every word in the garden path sentence: “The committee mentioned the issue would cause a problem at the meeting.” Note that in the last PCA plot, the red square on the left is the last token in the garden path sentence and the red square on the right is the last token in the disambiguated sentence.
Figure 9: Average movement per token and token movement distribution for disambiguated and garden path sentences.
Figure 10: Distribution of token stability across garden path and disambiguated sentences
Figure 11: Distribution of the inverse normalized token stability for garden path and disambiguated sentences
Figure 12: Distribution of the inverse normalized token stability for garden path and disambiguated sentences when normalizing by distance from the critical region in terms of word increments.
Figure 13: Summary statistics of token embedding distance between model interpretation of garden path and disambiguated sentences across RoBERTa, DeBERTa, and DistilBERT. y -axis on the left 2 plots is the cosine similarity between garden path sentences and disambiguated equivalents, y -axis on the right two plots is Euclidean distance between garden path sentences and disambiguated equivalents. Upper 2 plots are showing the average behavior of the CLS token and lower 2 plots are showing the average behavior of the hinge token.
Figure 14: From left to right: Average cosine similarity between the hinge token in garden path and disambiguated sentences as additional tokens are added to context, Average Euclidean distance between the hinge token in garden path and disambiguated sentences as additional tokens are added to context. Upper 2 are RoBERTa, lower 2 are OLMo.
Figure 15: Training and Validation accuracy curves for the word sense classifier
Figure 16: Classifier performance on increments calculated by different functions (from left to right) no perturbations, all characters lower case, and all characters upper case.
Figure 17: Classifier performance on increments as represented by f1 with (from left to right) no perturbations, all characters lower case, and all characters upper case.
Figure 18: Classifier performance on data generated by D4 with various perturbations including all characters lowercase, all characters upper case, and no punctuation
Figure 19: Classifier performance on data generated by D2 with various perturbations including all characters lowercase, all characters upper case, and no punctuation
Figure 20: Classifier performance on data generated by D3 with various perturbations including all characters lowercase, all characters upper case, and no punctuation
Figure 21: Classifier performance on data generated by D1 with various perturbations including all characters lowercase, all characters upper case, and no punctuation
Figure 22: Average trajectory of a garden path sentence through embedding space.
Figure 23: Average similarity statistics across all garden path sentences in the dataset when incrementally embedded into the OLMo embedding space. From left to right: Average cosine similarity between the hinge token in garden path and disambiguated sentences as additional tokens are added to context, Average Euclidean distance between the hinge token in garden path and disambiguated sentences as additional tokens are added to context.
Figure 24: UMAP Clustering of CLS token embeddings of “run” sentences. All imperatives were grouped into cluster 0, all interrogatives were grouped into cluster 2, declaratives were split across clusters 1 and 2, and exclamatory sentences were spread across all 3 clusters.
Figure 25: Principle Component Analysis of the latent space representation of the target token (“run”) of sentences, grouped by sentence type.
Figure 26: Semantic similarity metrics across the average CLS embedding trajectory of all garden path and negated garden path sentences. The plot on the left shows the cosine similarity between the two average trajectories and the plot of the right shows the Euclidean distance between the two average trajectories.
Figure 27: Plot depicting the number of times that each corresponding token is the most turbulent token in its respective sentence.
Figure 28: Training and Validation accuracy curves for the garden path variety classifier.
Figure 29: Training and validation accuracy curves of the classifier model trained on shuffled data.
Figure 30: PCA plot showing the characteristic gap between parses of a garden path sentence and the unambiguous equivalent.
Figure 31: PCA plot CLS token trajectory and ambiguous token trajectory. The left plot shows the CLS token trajectories and the right plot shows the ambiguous token trajectories.
Figure 32: Average cosine similarity and euclidean distance of the token that exhibits the most displacement over the trajectory of garden path sentences.
Figure 33: Principle Component Analysis of the latent space representation of the CLS token of sentences with the word run, grouped by sentence type.
Figure 34: Principle Component Analysis of the latent space representation of the CLS token of sentences with the word run, grouped by sentence type, with labelled cluster centroids.
Figure 35: Sequential PCA plots of CLS token embedding after adding 1 token to context per step showing emergence of organization in latent space.
Figure 36: Principle Component Analysis of the latent space representation of the incremental trajectory of sentences containing the token “run”.
Figure 37: Principle Component Analysis of the latent space representation of the target token (“run”) of sentences, grouped by sentence type.
Figure 38: Validation and training accuracy curves for classifier trained on the trajectory of the first token in each sentence.
Figure 39: Validation and training accuracy curves for fully-connected feed-forward classifier trained on the difference vector between the penultimate and last token in each sentence.
Figure 40: Validation and training accuracy curves for classifier trained on the trajectory of the first token in the critical region of each sentence.
Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence level by examining the per-token mean log-probability of a sentence instead of the probabilities of individual tokens. We investigate sentence-level contextual entrainment across 26 LLMs from seven families and two datasets, which cover both subjective and objective tasks. We find that sentence-level contextual entrainment exists. This means that the sentences in the prompt (even if they are counterfactual statements) can significantly increase their probability during model inference time. As the model size increases, contextual entrainment gradually decreases. We also find that contextual entrainment is controlled by 2% to 4% of the attention heads. Turning off these attention heads can effectively mitigate contextual entrainment without hurting the model's performance.
We introduce methods to quantify how Large Language Models (LLMs) encode and store contextual information, revealing that tokens often seen as minor (e.g., determiners, punctuation) carry surprisingly high context. Notably, removing these tokens -- especially stopwords, articles, and commas -- consistently degrades performance on MMLU and BABILong-4k, even if removing only irrelevant tokens. Our analysis also shows a strong correlation between contextualization and linearity, where linearity measures how closely the transformation from one layer's embeddings to the next can be approximated by a single linear mapping. These findings underscore the hidden importance of filler tokens in maintaining context. For further exploration, we present LLM-Microscope, an open-source toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions (via an adapted Logit Lens), and measures the intrinsic dimensionality of representations. This toolkit illuminates how seemingly trivial tokens can be critical for long-range understanding.
Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4
AIRI · Skoltech · Lomonosov Moscow State University +1
Human language comprehension unfolds sequentially: each word is processed in the context of those that came before, and the interpretation builds incrementally over time. Surprisal, the negative log probability of a word given its context, has been the dominant predictor of incremental processing cost. But surprisal reduces rich sequential representations to a single scalar at each word, discarding information about the direction in which the interpretation has been evolving. Dynamical-systems approaches suggest that the trajectory of the evolving interpretive state, not just its position at each moment,should shape processing, and language itself may have local momentum, since speakers plan utterances a few words at a time. We introduce trajectory extrapolation error: at each word, we fit a linear trajectory to the preceding hidden states of a transformer language model and measure deviation from the extrapolated path. On the Natural Stories corpus, this measure is nearly orthogonal to surprisal (r = .044) and independently predicts self-paced reading times. The effect is especially pronounced in garden-path sentences, strengthens with model scale (GPT-2 Small to Large), and replicates across architectures with different positional encoding schemes (GPT-2 vs. Pythia/RoPE). A displacement control shows the effect is not reducible to representational change magnitude: displacement and extrapolation error predict in opposite directions. These findings reveal two dissociable components of processing cost: word-level prediction error (surprisal) and sensitivity to the local momentum of the unfolding interpretation (trajectory extrapolation error).
Elan Barenholtz
Machine Perception & Cognitive Robotics Laboratory Department of Psychology / Center for Complex Systems Florida Atlantic University