Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.
Figures & tables
Fig. 1: Similar specificity scores may be supported by different structural evidence under cross-domain application. The matched score is illustrative, not a verified sentence-level measurement.
Domain
Sentences
Avg. Sent. Len
Wikipedia
900,406
21.96
GitHub Docs
170,798
15.24
Ansible Docs
44,452
11.73
Python Ref
179,549
11.64
TABLE I: Summary statistics for each corpus after preprocessing.
Fig. 2: Candidate selection as evaluated and as a possible documentation workflow. Left: three score-blind candidates are fixed before post-generation scoring, directional selection, and blinded review. Right: a proposed workflow could place the same ranking step after factual and semantic checks and before an author decision. Dashed boxes mark the proposed extension.
Gemma
GPT-OSS-120B
Policy
Direction-valid
Gain [95% CI]
Direction-valid
Gain [95% CI]
First candidate
43/60 (71.7%)
reference
31/60 (51.7%)
reference
SpeciTeller
50/60 (83.3%)
+11.7 [+1.7, +21.7]
34/60 (56.7%)
+5.0 [ −6.7 , +16.7]
Adapt-1
46/60 (76.7%)
+5.0 [ −3.3 , +13.3]
34/60 (56.7%)
+5.0 [ −6.7 , +16.7]
Adapt-2
45/60 (75.0%)
+3.3 [ −6.7 , +13.3]
34/60 (56.7%)
+5.0 [ −6.7 , +16.7]
Adapt-3
45/60 (75.0%)
+3.3 [ −6.7 , +13.3]
36/60 (60.0%)
+8.3 [ −1.7 , +20.0]
TABLE II: Post-generation ranking by generator and review session. Each cell reports direction-valid selections and the paired percentage-point gain over that generator’s first candidate, with 95% source-case bootstrap intervals. Adapt-1/2/3 repeat one pinned implementation; Adapt mean and consensus are secondary.
Metric
Wikipedia
GitHub Docs
Ansible Docs
Python Ref
Score level, variability, and agreement
SpeciTeller mean
.7534
.4960
.4257
.2921
Adapt mean
.3490
.3993
.3704
.2825
Adapt row SD
.0239
.0324
.0171
.1542
ST–Adapt mean ρ
.5103
.0704
−.0656
.4430
Related granularity measure
TABLE III: Cross-context measurement profile. Higher ST/Adapt means more specific; higher native GS means coarser. Means retain native scales; other entries are within-corpus Spearman correlations. Adapt mean is secondary; row SD summarizes variation across three repetitions.
Corpus
Keep
Mean
Median
(%)
Orig.
Strict
Orig.
Strict
Wikipedia
35.6
0.753
0.764
0.884
0.942
GitHub Docs
46.1
0.496
0.340
0.543
0.234
Ansible Docs
48.5
0.426
0.388
0.369
0.298
Python Ref
70.9
0.292
0.284
0.154
0.127
TABLE IV: Preprocessing sensitivity on original and strict subsets. “Keep” is the percentage retained by the strict condition.
Rows
Corpus
Raw
Standardized [95% CI]
Ret.
Orig.
GitHub
0.257
0.243 [0.236, 0.249]
94.5%
Orig.
Ansible
0.328
0.271 [0.253, 0.290]
82.6%
Orig.
Python
0.461
0.329 [0.303, 0.355]
71.3%
Strict
GitHub
0.425
0.304 [0.295, 0.313]
71.6%
Strict
Ansible
0.376
0.258 [0.236, 0.281]
68.5%
Strict
Python
0.481
0.236 [0.208, 0.264]
49.1%
TABLE V: Raw and token-length-standardized Wikipedia-minus-technical SpeciTeller gaps. Brackets give 95% document-cluster bootstrap intervals; “Ret.” is the percentage of raw gap magnitude retained.
TABLE VII: Unified pilot Spearman correlation ( n=40 per corpus). ST and comparator Adapt-1/2/3 are specificity predictors; these are repeated fits of the same pinned post-publication implementation. Qwen, Gemma, GPT-OSS-20B, and GPT-OSS-120B are matched LLM rubric judges. Pooled cells give paired-bootstrap 95% intervals. Adapt mean and GS-align are secondary; GS-align negates native GranuScore rank association and is not a raw-score conversion.
Corpus
Quintile
Mean score
Mean pooled label
Ansible
Q1
0.001
0.344
Ansible
Q2
0.314
0.620
Ansible
Q3
0.404
0.760
Ansible
Q4
0.574
0.798
Ansible
Q5
1.000
0.656
GitHub
Q1
0.001
0.260
Appendix
TABLE VIII: Binned correspondence between SpeciTeller scores and pooled pilot labels.
Feature
Wiki
GitHub
Ansible
Python
TFIDF-mean
-0.729
-0.387
-0.229
-0.167
TFIDF-max
-0.682
-0.289
-0.176
-0.092
Tech-ratio
0.359
0.424
0.374
0.334
Tokens
0.754
0.417
0.312
0.242
Chars
0.695
0.450
0.318
0.248
Appendix
TABLE IX: Spearman correlation between SpeciTeller and visible lexical cues.
TABLE XI: Controlled-edit robustness under within-corpus normalization.
Corpus
ST mean
Adapt mean
Adapt-1/2/3 means
Row SD
Spearman [95% CI]
Wikipedia
.7534
.3490
.3475/.3786/.3208
.0239
.5103 [.5008, .5191]
GitHub Docs
.4960
.3993
.3664/.4418/.3898
.0324
.0704 [.0588, .0825]
Ansible Docs
.4257
.3704
.3857/.3480/.3774
.0171
−.0656 [ −.1092 , −.0285 ]
Python Ref
.2921
.2825
.4220/.0680/.3574
.1542
.4430 [.4120, .4819]
Wikipedia-minus-technical native-scale gaps
Technical corpus
ST gap [95% CI]
Adapt mean gap [95% CI]
Appendix
TABLE XII: Original-row comparison of SpeciTeller (ST) and the pinned post-publication comparator. Adapt mean averages Adapt-1/2/3 by sentence; row SD is the mean sentence-level population SD across repetitions. Intervals resample documents.
Source/view
Corpus
Range
L5
U5
SpeciTeller
Wikipedia
.814
1.85
41.47
SpeciTeller
Python
.938
33.27
4.59
Adapt-1
Python
.351
0.00
0.02
Adapt-2
Python
.131
42.26
0.01
Adapt-3
Python
.360
0.00
0.02
GranuScore
GitHub
.399
0.08
1.55
Appendix
TABLE XIII: Selected output-distribution diagnostics. “Range” is (p95−p05)/(U−L) ; L5/U5 are percentages in the lower/upper outer 5%. GS-units removes the no-unit sentinel.
Corpus
Identifier type
Pinned identifier
GitHub Docs
Source Git commit
5b8a8c96f169
Ansible Docs
Source Git commit
83bbdd618bb6
Python Ref
Source archive SHA-256
9e98419a01ef
Wikipedia
Source archive SHA-256
5b08ca86a17f
Wikipedia
Deterministic build
2,000 files
Appendix
TABLE XIV: Corpus source provenance. Supporting reproducibility files are in the SpecTech repository.
TABLE XVII: Experimental roles and inputs. Rubric-judge results concern the 80-sentence pilot, not either generated candidate set; a model name appearing in both rows does not imply self-judging.