cs.CLOct 4, 2026
SaveWriting as a Self-Organized Critical Process
Organizations: NTR Labs / Moscow, Russia · Higher IT School of Tomsk State University / Tomsk, Russia
Abstract
We explain autocorrelation decay power laws omnipresent in texts by self-organized criticality. Specifically, we analyze the recently released KLiCKe keystroke dataset and show that not only the final texts' autocorrelations form a manifold that adheres to a power law with a finite-size scaling, but also the text revisions generate revision-size-dependent restoring dynamics toward that manifold. Thus, human writing appears to dynamically regulate semantic correlations in a text toward a critical state.
Figures & tables
Figure 1: Finite-size collapse of normalized semantic autocorrelations in completed KLiCKe essays. (a) Normalized excess autocorrelation for bins of final text length. (b) Data-collapse score as a function of the horizontal rescaling exponent , with the optimum at . (c) The same autocorrelation curves after rescaling the lag as , showing improved collapse across text lengths.
| Measure | Range / threshold | Exponent |
|---|---|---|
| Burst size | ||
| Word loss | ||
| Revision size | – | |
| Burst duration |
Table 1: Tail behavior across writing-activity measures. Size measures use different units: events ( ), net words ( ), and gross characters ( ). The revision entry summarizes fits across lower thresholds of 100–500 characters. Duration entries are density-equivalent exponents derived from CCDF slopes below/above , rather than a single fitted power-law density.
Figure 2: Emergence of long-range semantic correlations during writing. Curves show mean excess FastText autocorrelation at five stages of normalized writing time for a fixed cohort of 612 essays with final lengths of 250–349 words. The same essays contribute at every stage and plotted octave; error bars denote 95% confidence intervals.
Figure 3: Illustrative writing trajectory exhibiting approach, persistent capture, and confinement near final-text reference tube.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Estimator changes and net restoration across revision sizes. Results use the common 4,241-essay cohort and retain bulk edits. (A) Fraction of revisions changing the estimator’s normalized token sequence. (B) Event-weighted mean , corrected by subtracting the calibrated pre-minus-post error-variance difference ( , 16-word blocks). Positive values indicate average movement toward the length-conditioned reference. Bars are pointwise 95% whole-essay bootstrap intervals; bin counts are shown. The horizontal axis is symmetric logarithmic, linear within , to display both the small central-bin effects and the uncertain extreme-bin estimates. These marginal means are distinct from covariate-adjusted feedback slopes.
| Essay | Observed editing pattern | ||||
|---|---|---|---|---|---|
| 10768319 | 4 | 1,439 | 1 | 0 | Repetitive nonlexical strings inserted over several episodes, then removed in bulk. |
| 68216667 | 4 | 1,095 | 1,095 | 42.10 | Extended insertion and character-by-character removal of repetitive letter strings. |
| 23987928 | 4 | 518 | 332 | 133.19 | Insertion, replacement, and deletion of nonlexical keyboard sequences. |
| 30578676 | 3 | 836 | 1 | 0 | Nonlexical strings typed and removed, including a 471-character insertion followed by its bulk deletion. |
| 40319803 | 1 | 383 | 383 | 102.10 | Three-character deletion followed by 380 characters of on-topic prose. |
| 76169954 | 1 | 323 | 323 | 95.92 | Three-character deletion followed by 320 characters of on-topic prose. |
Table 2: Selected essay audits. counts complete extracted revisions exceeding 300 gross changed characters in each log. , , and describe the largest such revision: gross characters, native modifications, and duration in seconds. Descriptions summarize the inspected large episodes, not necessarily only the largest. The cases are selected examples, not an exhaustive ranking or a representative sample.
Figure 5: Temporal avalanches and deletion runs. (A) Gap-based avalanche definition. (B) Empirical size CCDF and discrete power-law tail above , with probability-mass exponent 3.159. (C) Cadence-normalized duration CCDF; the descriptive two-regime fit has crossover , with the shaded 95% whole-essay bootstrap interval . The dashed gray curve is the rejected single-regime fit. (D) Net word loss per deletion run, conditional on positive loss; the fitted tail begins at six words and has probability-mass exponent 2.640. The original saved panels are reproduced without refitting.
Figure 6: Revision and production size distributions. (A,B) Empirical survival probabilities conditional on and four fitted distribution families. The additional production power law above 300 characters is plotted only on its support and multiplied by to share the conditional-on-100 probability scale. The guide at 300 is not an estimated breakpoint. Points use the saved subsampled empirical plotting grid. (C) Power-law exponents across fixed lower thresholds, with pointwise 95% whole-essay bootstrap intervals from refitting. Bulk edits remain included throughout.
| Episode type | Episodes | Essays | [95% interval] | |
|---|---|---|---|---|
| Revision | 100 | 5,352 | 2,526 | 2.995 [2.927, 3.062] |
| Revision | 200 | 1,425 | 935 | 3.127 [2.995, 3.281] |
| Revision | 300 | 591 | 410 | 3.170 [2.980, 3.429] |
| Revision | 400 | 283 | 196 | 2.856 [2.638, 3.143] |
| Revision | 500 | 183 | 130 | 2.784 [2.556, 3.099] |
| Production | 100 | 12,742 | 3,709 | 4.174 [4.110, 4.243] |
Table 3: Fixed-threshold gross-character fits, including bulk edits. Intervals are pointwise 95% whole-essay bootstrap intervals; all ten rows had 500 valid refits. Counts at successive thresholds are nested and must not be added.
| Episode type | Test episodes | Alternative | [95% interval] | |
|---|---|---|---|---|
| Production | 400 | 36 | Exponential | -0.1117 [-0.2409, 0.0277] |
| Production | 400 | 36 | Lognormal | -0.0299 [-0.0495, -0.0082] |
| Production | 400 | 36 | Power law with cutoff | -0.0323 [-0.0529, -0.0086] |
| Revision | 100 | 1,598 | Exponential | 0.1942 [0.1152, 0.2835] |
| Revision | 100 | 1,598 | Lognormal | 0.0003 [-0.0006, 0.0011] |
| Revision | 100 | 1,598 | Power law with cutoff | 0.0006 [-0.0005, 0.0018] |
Table 4: Held-out comparisons for gross character volume. is mean test log likelihood per episode for the power law minus that for the named alternative; positive values favor the power law. Intervals resample test essays conditional on trained models. Only bulk-included primary fits are shown.
Explore similar work
We propose a statistical-field framework for text generated by large language models (LLMs), treating token embeddings as continuous spin variables on a one-dimensional chain. Defining a susceptibility from the connected two-point correlator and an order parameter from the ensemble-averaged embedding field, we vary the \texttt{softmax} temperature and observe a sharp susceptibility peak near a characteristic with power-law-like scaling, a concurrent rapid change in the order parameter, and a collapse onto a single semantic direction below . The intrinsic dimension estimated by the two nearest neighbor (TwoNN) method independently corroborates these findings, reaching a minimum near . Results are robust across model scales (Qwen3: 0.6B--32B) and prompt categories. While the phenomenology closely resembles a continuous phase transition, the non-equilibrium nature of autoregressive generation warrants further investigation. Our framework provides quantitative tools for probing the collective statistical structure of LLM outputs and suggests connections between decoding strategies and critical phenomena.
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
Successive self-training on a language model's own outputs is widely characterized as a process of flattening: diversity drops, distributions narrow, and the text becomes "more like itself." We provide evidence that this characterization is incomplete. Across eleven generations of self-training on five models (GPT-2 124M, Pythia-410M, Pythia-1.4B, OPT-1.3B, Pythia-2.8B), language is not flattened uniformly -- it is restructured. Surface markers (discourse connectives, hedges, em-dashes) rise, while mid- and deep-syntactic structures (questions, parentheticals, passives, subjunctives) collapse. We formalize this asymmetric collapse as the Structural Depth Hypothesis (SDH): the per-generation decay rate of a linguistic feature is predicted primarily by its structural depth -- the number of nested syntactic dependencies it requires -- and only secondarily by its generation-zero output frequency. Pooling 17-feature panels from five models spanning three architecture families (N=85), the pooled Spearman correlation is rho=0.540 (p < 10^{-6}; cluster-bootstrap 95% CI [0.434, 0.634]), while frequency is a substantially weaker predictor (rho=0.225). A matched human-text fine-tuning control yields rho=0.039 (p=0.88), confirming the gradient is self-training-specific. We further document a Superficial Complexity Paradox: aggregate complexity proxies (dep-tree depth, TTR, word length) all rise as the underlying clause structure dies, with direct implications for training-data curation and LLM-text detection.
Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing
Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produced it. Final text alone cannot reveal whether a document was produced through human typing, AI generation, or mixed human-AI collaboration. Existing process-tracking tools help, but many are tied to host-document histories, provide coarse activity records, and offer limited control over the writing environment. Humanly is a writing platform that makes the writing process itself the evidence. Users configure writing environments for personal documents or assigned tasks and draft in a workspace that records writing activity and in-platform AI assistance. Humanly can package a completed session into a sealed writing certificate with configuration-aware anomaly behavior review. It can support writing scenarios such as course assignments, peer review, and personal certification. Our user study shows that Humanly is helpful across roles, and a red-teaming study shows that the Humanly Typing Detector distinguishes human hand typing from automated typing.