Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor
Organizations: University of Central Florida Orlando, Florida, USA
Abstract
Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.
Figures & tables
| Family | Method | Floor | In-span | Ratio | ||
|---|---|---|---|---|---|---|
| Antisense | R1 GNN | 14 | 0.018972 | 0.030885 | 0.049856 | 2.627909 |
| Antisense | R1 no-message | 14 | 0.018972 | 0.031574 | 0.050546 | 2.664273 |
| Antisense | Chemistry tree | 14 | 0.018972 | 0.023931 | 0.042903 | 2.261392 |
| Antisense | Token CNN | 14 | 0.018972 | 0.030339 | 0.049311 | 2.599177 |
| Sense | R1 GNN | 10 | 0.003320 | 0.003335 | 225.665139 | |
| Sense | R1 no-message | 10 | 0.003221 | 0.003236 | 218.972986 |
| Grouped benchmark: 1,924 of 2,927 rows (65.73%) from one source in one sequence component | |||||||
|---|---|---|---|---|---|---|---|
| Sensitivity | Method | Grouped MSE | APP MSE | S7 MSE | |||
| Source-excluded | R1 GNN | 0.0641 | -0.212 | 0.1631 | -0.005 | 0.1191 | -0.324 |
| R1 no-message | 0.0625 | -0.182 | 0.1637 | -0.009 | 0.1221 | -0.358 | |
| Chemistry tree | 0.0529 | -0.001 | 0.1655 | -0.020 | 0.0974 | -0.082 | |
| Token CNN | 0.0628 | -0.187 | 0.1626 | -0.002 | 0.1103 | -0.226 | |
| Training row mean | 0.0571 | -0.079 | 0.1810 | -0.115 | 0.1587 | -0.764 | |
| Panel/model | Inputs | Forced | AS | SS | NA | Pred. NA | ||
|---|---|---|---|---|---|---|---|---|
| Measured panel | 165 | – | 0.079894 | – | 0.041341 | 0.020415 | 0.018138 | – |
| R1 GNN | 90 | 0 | 0.077350 | 0.000041 | 0.042930 | 0.016196 | 0.018183 | 0.000002 |
| R1 no-message | 90 | 0 | 0.077353 | 0.000005 | 0.042811 | 0.016371 | 0.018166 | 0.000002 |
| Chemistry tree | 90 | 0 | 0.078546 | 0.000012 | 0.043859 | 0.016079 | 0.018596 | 0.000146 |
| Token CNN | 90 | 0 | 0.077276 | 0.000211 | 0.042819 | 0.016072 | 0.018174 | 0.000005 |
| Quantity | Model MSE | Matched control | Difference | Pearson | |
|---|---|---|---|---|---|
| Endpoints, released test split | 6078 | 0.084224 | 0.259250 | -0.175026 | 0.822247 |
| Endpoints, rectangle population | 3696 | 0.089748 | 0.246508 | -0.156759 | 0.797282 |
| Single substitutions | 12486 | 0.155262 | 0.302181 | -0.146919 | 0.697187 |
| Interactions (primary) | 2083 | 0.361690 | 0.456701 | -0.095011 | 0.462959 |
| Interactions, library 2 | 1947 | 0.327946 | 0.455823 | -0.127877 | 0.531564 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Inhibition fraction, by cohort: the grouped benchmark uses the released inhibition fraction; APP uses ; the primary B3 panel uses (Appendix A5.1 ); Davis S7 is kept on its original response scale. No clipping. | |
| Population mean of that assay endpoint; not directly known from a reported sample mean. | |
| Committed prediction of model from the permitted molecule/context features . | |
| Study/source-linked component and sequence-only component, respectively. | |
| Within-partition observation weight , where is the number of study components and is its size. | |
| Node features, relation adjacency tensor, node mask, and training-masked numeric context; here is the context tensor, not the contrast matrix of Section 2 . |
| Primary subscript | Frozen name |
|---|---|
| OMe | 2-O-Methyl |
| F | 2-Fluoro |
| DNA | 2-Deoxy |
| AEM | 2-Aminoethoxymethyl |
| APM | 2-Aminopropoxymethyl |
| EA | 2-Aminoethyl |
| Reconciliation diagnostic | Count/value |
|---|---|
| identity_matched_pairs | 1604 |
| within_half_tenth_percentage_point | 1209 |
| within_half_percentage_point | 1258 |
| greater_than_five_percentage_points | 299 |
| max_absolute_activity_difference | 1.0852170324990895 |
| Cohort | Model | MSE | Pred. SD | Bias | ||
|---|---|---|---|---|---|---|
| APP | R1 GNN | 0.1625 | -0.001 | -0.018 | 0.0072 | -0.0081 |
| R1 no-message | 0.1630 | -0.005 | 0.015 | 0.0085 | -0.0286 | |
| R0 GNN | 0.1625 | -0.001 | -0.026 | 0.0073 | -0.0049 | |
| R0 no-message | 0.1629 | -0.004 | -0.001 | 0.0085 | -0.0241 | |
| Chemistry tree | 0.1896 | -0.168 | -0.021 | 0.0103 | -0.1644 | |
| Token CNN | 0.1644 | -0.013 | 0.080 | 0.0036 | -0.0483 |
| Quantity and scope | Actual value |
|---|---|
| Core fit objects | 2312 |
| Executed optimizer updates | 1407967 |
| Selected-checkpoint updates | 406367 |
| Summed core-fit CPU seconds | 12523.769 |
| GPU training-region elapsed seconds | 11689.879 |
| Completed stage child CPU seconds | 13202.199 |
| Sensitivity | Method | Model MSE | Zero MSE | Flips | ||
|---|---|---|---|---|---|---|
| Omit AS | Tree | 154/130 | 0.059628–0.077251 | 0.056062–0.073257 | 2.456–4.031 | 0 |
| Omit AS | GNN | 154/130 | 0.056195–0.073399 | 0.056062–0.073257 | 0.015–0.143 | 0 |
| Omit AS | No-msg | 154/130 | 0.056160–0.073339 | 0.056062–0.073257 | -0.045–0.098 | 1 |
| Omit AS | CNN | 154/130 | 0.056290–0.073475 | 0.056062–0.073257 | 0.007–0.229 | 0 |
| Omit SS | Tree | 150/126 | 0.055818–0.079524 | 0.053743–0.075358 | 1.176–4.188 | 0 |
| Omit SS | GNN | 150/126 | 0.053882–0.075507 | 0.053743–0.075358 | -0.009–0.151 | 1 |
| Sensitivity | Method | Predicted SD | Recorded SD | Largest influence | Gap change |
|---|---|---|---|---|---|
| Omit AS | Tree | 0.008806–0.018376 | 0.204264–0.229500 | JC10 | -0.001278 |
| Omit AS | GNN | 0.000252–0.001648 | 0.204264–0.229500 | JC10 | -0.000115 |
| Omit AS | No-msg | 0.000299–0.001714 | 0.204264–0.229500 | JC10 | -0.000122 |
| Omit AS | CNN | 0.001113–0.002984 | 0.204264–0.229500 | JC10 | -0.000197 |
| Omit SS | Tree | 0.007219–0.018653 | 0.203193–0.230994 | JC5 | -0.002558 |
| Omit SS | GNN | 0.000327–0.001674 | 0.203193–0.230994 | JC5 | -0.000139 |
| Archive path | Preserved evidence |
| v4/adjudication/row_reconciliation.csv | All 1,972 raw source rows (1,924 admitted): cells, chemistry, cardinality, original/primary values, categories |
| v4/dataset/primary_reanchored_observations.jsonl | 2,607 versioned observations, original values and graph indices |
| v4/review_checks/source_scores.csv | All-source, Bramsen, other-source, every individual source and constant; three weightings |
| v4/evaluation/per_source_metrics.csv | Every source after focused refitting; exact matched coverage |
| v4/evaluation/ensemble_predictions.csv | Exact member means and cohort identifiers |
| v4/fits/ | Best/last checkpoints, optimizer states, histories, selected transforms and memberships |
| Cohort | Original v3 model | MSE | Pred. SD | Bias | ||
|---|---|---|---|---|---|---|
| APP | R1 GNN | 0.1625 | -0.001 | -0.018 | 0.0072 | -0.0081 |
| R1 no-message | 0.1630 | -0.005 | 0.015 | 0.0085 | -0.0286 | |
| Chemistry tree | 0.1896 | -0.168 | -0.021 | 0.0103 | -0.1644 | |
| Token CNN | 0.1644 | -0.013 | 0.080 | 0.0036 | -0.0483 | |
| S7 | R1 GNN | 0.1093 | -0.215 | 0.112 | 0.0059 | 0.1403 |
| R1 no-message | 0.1047 | -0.164 | 0.117 | 0.0066 | 0.1231 |
| Method | MSE | Pred. SD | Label SD | Retro. mean | Train const | ||
|---|---|---|---|---|---|---|---|
| GNN | 0.077350 | 0.031844 | 0.184788 | 0.040265 | 0.282656 | 0.079894 | 0.080173 |
| No-msg | 0.077353 | 0.031810 | 0.185135 | 0.038465 | 0.282656 | 0.079894 | 0.080173 |
| Tree | 0.078546 | 0.016873 | 0.167266 | 0.076857 | 0.282656 | 0.079894 | 0.080173 |
| CNN | 0.077276 | 0.032767 | 0.190711 | 0.045136 | 0.282656 | 0.079894 | 0.080173 |
| Method | Family | MSE | Zero MSE | Pred. SD | Label SD | ||
|---|---|---|---|---|---|---|---|
| GNN | AS effect | 14 | 0.049856 | 0.051986 | 0.028779 | 0.024352 | 0.170616 |
| GNN | SS effect | 10 | 0.003335 | 0.000855 | 0.286065 | 0.034469 | 0.013953 |
| GNN | Interaction | 140 | 0.068335 | 0.068205 | -0.090263 | 0.001589 | 0.222442 |
| No-msg | AS effect | 14 | 0.050546 | 0.051986 | 0.019753 | 0.021060 | 0.170616 |
| No-msg | SS effect | 10 | 0.003236 | 0.000855 | 0.269464 | 0.034474 | 0.013953 |
| No-msg | Interaction | 140 | 0.068281 | 0.068205 | -0.034429 | 0.001652 | 0.222442 |
| Strand | Indistinguishable variants | Positions | Size |
|---|---|---|---|
| AS | JC-A1, JC-F1, JC-S1 | 3, 18 | 3 |
| AS | JC-A2, JC-F2, JC-S2 | 4, 18 | 3 |
| AS | JC-A3, JC-F3, JC-S3 | 3, 4, 18 | 3 |
| SS | DO003, DO004 | 17 | 2 |
| Scen. | Weights | Model | Constant | Difference | Lower | Upper | Grp |
|---|---|---|---|---|---|---|---|
| primary | rows | GNN | row const | -0.001217 | -0.004579 | 0.020462 | 10 |
| primary | rows | GNN | grp const | -0.006226 | -0.010077 | 0.021086 | 10 |
| primary | rows | Tree | row const | -0.006784 | -0.013080 | 0.003109 | 10 |
| primary | rows | Tree | grp const | -0.011792 | -0.016886 | 0.000805 | 10 |
| primary | equal study | GNN | row const | 0.015532 | -0.007321 | 0.038432 | 10 |
| primary | equal study | GNN | grp const | 0.015232 | -0.007620 | 0.038511 | 10 |
| Method | Seed | Scale | Changed | Mean ovl. | Min ovl. | Med. disp. | Max disp. |
|---|---|---|---|---|---|---|---|
| GNN | 1103 | -0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| GNN | 1103 | 0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| GNN | 2207 | -0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| GNN | 2207 | 0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| GNN | 3301 | -0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| GNN | 3301 | 0.250000 | 0 | 1.000000 | 1.000000 | 0.000000 | 0.000000 |
| Arm | Seed | Scale | Max coordinate | Max | Max relative |
|---|---|---|---|---|---|
| R1 | 1103 | -0.250000 | 0.000000 | 0.000000 | 0.000000 |
| R1 | 1103 | 0.250000 | 0.000000 | 0.000000 | 0.000000 |
| R1 | 2207 | -0.250000 | 0.000000 | 0.000000 | 0.000000 |
| R1 | 2207 | 0.250000 | 0.000000 | 0.000000 | 0.000000 |
| R1 | 3301 | -0.250000 | 0.000000 | 0.000000 | 0.000000 |
| R1 | 3301 | 0.250000 | 0.000000 | 0.000000 | 0.000000 |
| Protocol | GNN | No-message | Tree | CNN |
|---|---|---|---|---|
| v1 original | 0.521 | 0.497 | 0.565 | 0.453 |
| v2 factorial | 0.538 | 0.498 | 0.574 | 0.530 |
| v2 deployment | 0.397 | 0.367 | 0.631 | 0.510 |
| v3 group selected | 0.534 | 0.577 | 0.445 | 0.583 |
| Source-excluded | 0.400 | 0.374 | 0.696 | 0.421 |
| Primary-assay | 0.565 | 0.588 | 0.514 | 0.563 |
| Method | Row MSE | Comp. MSE | Pred. SD | Sign | |
|---|---|---|---|---|---|
| Activity GNN | 0.2351 | 0.3560 | 0.083 | 0.036 | 111/137 |
| Pair GNN | 0.2221 | 0.3211 | 0.248 | 0.038 | 111/137 |
| Pair no-message | 0.2220 | 0.3196 | 0.228 | 0.043 | 111/137 |
| Pair tree | 0.2438 | 0.3101 | 0.120 | 0.167 | 109/137 |
| Pair ridge | 0.2193 | 0.2888 | 0.226 | 0.148 | 105/137 |
| Zero | 0.2808 | 0.4342 | – | 0.000 | 0/137 |
| Retention score | Model MSE | Zero MSE | Excess MSE |
|---|---|---|---|
| Support absence | 0.434741 | 0.434166 | +0.000575 |
| Probe sensitivity | 0.356843 | 0.356330 | +0.000513 |
| Ensemble SD | 0.363586 | 0.362827 | +0.000759 |
| Chemical novelty | 0.434741 | 0.434166 | +0.000575 |
| Random | 0.451729 | 0.451135 | +0.000594 |
| Coverage | Pair mass | Probe MSE | Its zero MSE | SD MSE | 95% interval | |
|---|---|---|---|---|---|---|
| 20% | 31.2 | 0.148697 | 0.150237 | 0.408344 | -0.002887 | [-0.004963,-0.000467] |
| 40% | 62.4 | 0.273094 | 0.273292 | 0.425092 | -0.001938 | [-0.003812,-0.000147] |
| 60% | 93.6 | 0.356843 | 0.356330 | 0.363586 | -0.000246 | [-0.001059,+0.000561] |
| 80% | 124.8 | 0.336510 | 0.336282 | 0.322389 | -0.000208 | [-0.000668,+0.000103] |
| 100% | 156.0 | 0.434741 | 0.434166 | 0.434741 | +0.000000 | [+0.000000,+0.000000] |
| Panel | Floor | Energy | Share | ||||||
|---|---|---|---|---|---|---|---|---|---|
| B3 (reproduced) | 165 | 90 | 140 | 140 | 72 | 68 | 0.009980 | 0.068205 | 14.63% |
| Bramsen, expanded | 1932 | 1183 | 1845 | 1845 | 1115 | 730 | 0.007565 | 0.055180 | 13.71% |
| new contrasts only | 1792 | 1111 | 1705 | 1705 | 1043 | 662 | 0.007367 | 0.054111 | 13.61% |
| MAVE-NN splicing | 3696 | 3696 | 2083 | 1986 | 1986 | 0 | 0 | 0.412839 | 0% |
| GB1 binding | 17689 | 17689 | 15856 | 15856 | 15856 | 0 | 0 | 0.549741 | 0% |