Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
Organizations: Mississippi State University, USA
Abstract
Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicated. This work audits sentence-specificity scoring artifacts on technical documentation and tests whether scores applied only after generation help choose among fixed LLM-generated revisions. Across Wikipedia and three technical-documentation corpora, the fixed general-domain predictor SpeciTeller and the pinned post-publication author-repository implementation of Ko et al.'s target-adapted predictor produce different corpus orders and same-sentence rank agreement from -0.066 to 0.510. Strict filtering and token-length adjustment change these patterns without reconciling them. In the Gemma set, SpeciTeller ranking raises direction-valid selection from 71.7% to 83.3% (+11.7 points; 95% source-case bootstrap interval +1.7 to +21.7); in the GPT-OSS-120B set, SpeciTeller ranking raises direction-valid selection from 51.7% to 56.7% (+5.0 points; 95% source-case bootstrap interval -6.7 to +16.7), and every primary single-score GPT-OSS-120B interval includes zero. These findings tie score interpretation and decision value to the predictor and candidate set.
Figures & tables
| Domain | Sentences | Avg. Sent. Len |
|---|---|---|
| Wikipedia | 900,406 | 21.96 |
| GitHub Docs | 170,798 | 15.24 |
| Ansible Docs | 44,452 | 11.73 |
| Python Ref | 179,549 | 11.64 |
| Gemma | GPT-OSS-120B | |||
|---|---|---|---|---|
| Policy | Direction-valid | Gain [95% CI] | Direction-valid | Gain [95% CI] |
| First candidate | 43/60 (71.7%) | reference | 31/60 (51.7%) | reference |
| SpeciTeller | 50/60 (83.3%) | +11.7 [+1.7, +21.7] | 34/60 (56.7%) | +5.0 [ , +16.7] |
| Adapt-1 | 46/60 (76.7%) | +5.0 [ , +13.3] | 34/60 (56.7%) | +5.0 [ , +16.7] |
| Adapt-2 | 45/60 (75.0%) | +3.3 [ , +13.3] | 34/60 (56.7%) | +5.0 [ , +16.7] |
| Adapt-3 | 45/60 (75.0%) | +3.3 [ , +13.3] | 36/60 (60.0%) | +8.3 [ , +20.0] |
| Metric | Wikipedia | GitHub Docs | Ansible Docs | Python Ref |
|---|---|---|---|---|
| Score level, variability, and agreement | ||||
| SpeciTeller mean | .7534 | .4960 | .4257 | .2921 |
| Adapt mean | .3490 | .3993 | .3704 | .2825 |
| Adapt row SD | .0239 | .0324 | .0171 | .1542 |
| ST–Adapt mean | .5103 | .0704 | .4430 | |
| Related granularity measure | ||||
| Corpus | Keep | Mean | Median | ||
|---|---|---|---|---|---|
| (%) | Orig. | Strict | Orig. | Strict | |
| Wikipedia | 35.6 | 0.753 | 0.764 | 0.884 | 0.942 |
| GitHub Docs | 46.1 | 0.496 | 0.340 | 0.543 | 0.234 |
| Ansible Docs | 48.5 | 0.426 | 0.388 | 0.369 | 0.298 |
| Python Ref | 70.9 | 0.292 | 0.284 | 0.154 | 0.127 |
| Rows | Corpus | Raw | Standardized [95% CI] | Ret. |
|---|---|---|---|---|
| Orig. | GitHub | 0.257 | 0.243 [0.236, 0.249] | 94.5% |
| Orig. | Ansible | 0.328 | 0.271 [0.253, 0.290] | 82.6% |
| Orig. | Python | 0.461 | 0.329 [0.303, 0.355] | 71.3% |
| Strict | GitHub | 0.425 | 0.304 [0.295, 0.313] | 71.6% |
| Strict | Ansible | 0.376 | 0.258 [0.236, 0.281] | 68.5% |
| Strict | Python | 0.481 | 0.236 [0.208, 0.264] | 49.1% |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Corpus | Edit | Mean | Med. | Cons. |
|---|---|---|---|---|
| GitHub | add_spec | 0.124 | 0.036 | 100% |
| GitHub | de_spec | -0.180 | -0.162 | 100% |
| GitHub | irrelev. | -0.001 | -0.003 | 80% |
| Ansible | add_spec | 0.054 | 0.020 | 75% |
| Ansible | de_spec | -0.161 | -0.023 | 91.7% |
| Ansible | irrelev. | -0.103 | -0.005 | 80% |
| Corpus | Score instance | Ann. A | Ann. B | Ann. C | Pooled [95% CI] |
|---|---|---|---|---|---|
| Ansible | ST | .188 | .502 | .733 | .480 [.126, .723] |
| Adapt-1 | .302 | .254 | .604 | .484 [.189, .686] | |
| Adapt-2 | .304 | .256 | .600 | .486 [.194, .688] | |
| Adapt-3 | .305 | .253 | .615 | .486 [.183, .696] | |
| Adapt mean | .304 | .249 | .600 | .480 [.184, .686] | |
| Qwen | .408 | .578 | .656 | .653 [.356, .853] |
| Corpus | Quintile | Mean score | Mean pooled label |
|---|---|---|---|
| Ansible | Q1 | 0.001 | 0.344 |
| Ansible | Q2 | 0.314 | 0.620 |
| Ansible | Q3 | 0.404 | 0.760 |
| Ansible | Q4 | 0.574 | 0.798 |
| Ansible | Q5 | 1.000 | 0.656 |
| GitHub | Q1 | 0.001 | 0.260 |
| Feature | Wiki | GitHub | Ansible | Python |
|---|---|---|---|---|
| TFIDF-mean | -0.729 | -0.387 | -0.229 | -0.167 |
| TFIDF-max | -0.682 | -0.289 | -0.176 | -0.092 |
| Tech-ratio | 0.359 | 0.424 | 0.374 | 0.334 |
| Tokens | 0.754 | 0.417 | 0.312 | 0.242 |
| Chars | 0.695 | 0.450 | 0.318 | 0.248 |
| Corpus | Tech-ratio | Strongest added | Indicator |
|---|---|---|---|
| GitHub | 0.424 | 0.392 | identifier |
| Ansible | 0.374 | 0.326 | version/numeric |
| Python Ref | 0.334 | 0.396 | version/numeric |
| Corpus | Edit | Raw mean | mean | Cons. | |
|---|---|---|---|---|---|
| GitHub | add_spec | 10 | 0.124 | 0.383 | 100% |
| GitHub | de_spec | 10 | -0.180 | -0.556 | 100% |
| GitHub | irrelev. | 10 | -0.001 | -0.004 | 80% |
| Ansible | add_spec | 8 | 0.054 | 0.178 | 75% |
| Ansible | de_spec | 12 | -0.161 | -0.530 | 91.7% |
| Ansible | irrelev. | 10 | -0.103 | -0.341 | 80% |
| Corpus | ST mean | Adapt mean | Adapt-1/2/3 means | Row SD | Spearman [95% CI] |
|---|---|---|---|---|---|
| Wikipedia | .7534 | .3490 | .3475/.3786/.3208 | .0239 | .5103 [.5008, .5191] |
| GitHub Docs | .4960 | .3993 | .3664/.4418/.3898 | .0324 | .0704 [.0588, .0825] |
| Ansible Docs | .4257 | .3704 | .3857/.3480/.3774 | .0171 | [ , ] |
| Python Ref | .2921 | .2825 | .4220/.0680/.3574 | .1542 | .4430 [.4120, .4819] |
| Wikipedia-minus-technical native-scale gaps | |||||
| Technical corpus | ST gap [95% CI] | Adapt mean gap [95% CI] | |||
| Source/view | Corpus | Range | L5 | U5 |
|---|---|---|---|---|
| SpeciTeller | Wikipedia | .814 | 1.85 | 41.47 |
| SpeciTeller | Python | .938 | 33.27 | 4.59 |
| Adapt-1 | Python | .351 | 0.00 | 0.02 |
| Adapt-2 | Python | .131 | 42.26 | 0.01 |
| Adapt-3 | Python | .360 | 0.00 | 0.02 |
| GranuScore | GitHub | .399 | 0.08 | 1.55 |
| Corpus | Identifier type | Pinned identifier |
|---|---|---|
| GitHub Docs | Source Git commit | 5b8a8c96f169 |
| Ansible Docs | Source Git commit | 83bbdd618bb6 |
| Python Ref | Source archive SHA-256 | 9e98419a01ef |
| Wikipedia | Source archive SHA-256 | 5b08ca86a17f |
| Wikipedia | Deterministic build | 2,000 files |
| Component | Identifier type | Pinned identifier |
|---|---|---|
| SpeciTeller | Repository Git commit | 218cb5a389b3 |
| SpeciTeller | Released-data SHA-256 | dd70443a77b6 |
| SpeciTeller | LIBLINEAR Git commit | 491c9f1188b9 |
| SpeciTeller | Container image tag | speciteller:py27 |
| Ko comparator | Upstream Git commit | 36f8e835e9dc |
| Ko comparator | Container image tag | spectech-ko-specificity:36f8e835-official-release-v1 |
| Component | Identifier type | Pinned identifier |
|---|---|---|
| Generator prompt G | Config-file SHA-256 | c14fc73dbf12 |
| Rubric prompt R | Config-file SHA-256 | 41686f9c2a98 |
| Ko released-domain diagnostic | Config-file SHA-256 | 9ea05db72f47 |
| Gemma review | Blinded review file SHA-256 | 73312e07427d |
| GPT-OSS-120B review | Blinded review file SHA-256 | 5c84775834cf |
| DGX run archive | Archive SHA-256 | bf6f5d9b5564 |
| Model(s) | Role | Visible input | Output | ID |
|---|---|---|---|---|
| gemma4:12b ; gpt-oss:120b | Candidate generator | Original sentence plus one add, reduce, or neutral request | One edited sentence | G |
| qwen3:14b ; gemma4:12b ; gpt-oss:20b ; gpt-oss:120b | Rubric judge | One pilot sentence | One integer, 1–5 | R |
| SpeciTeller; Adapt-1/2/3; GranuScore | Post-generation scorer | Fixed candidate text | Scalar score | — |