In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model runs the first layers of one model, converts the resulting intermediate representation once with the adapter, and then runs the remaining layers of the other model. Once the adapter is trained, several composed models that connect at different layers are obtained without retraining. Using the recurrent RWKV-4-Raven-7B and the Transformer-based Tulu-Pythia-6.9b, abbreviated as RWKV and Pythia, this study examines whether frozen models from different families can be recombined post hoc. The composed models answered multiple-choice questions, and those whose generations we examined produced syntactically well-formed text. The configuration that combines the first 5 layers of Pythia with the remaining 27 layers of RWKV reduced the Transformer key-value (KV) cache by 84.4% with accuracy not significantly different from that of RWKV alone. In multiple-choice accuracy, however, no composed model matched the parent model Pythia, and language-modeling performance decreased sharply on WikiText, a corpus of Wikipedia articles outside the training domain. The correspondence between intermediate representations was also obtained in one favorable case, with a shared tokenizer, the same depth, and the same hidden width, and does not show that the models share a general semantic space.
Figures & tables
Fig. 1: Shared-latent adapter. Solid lines are the self-reconstructions and the alignment that the loss includes, and dashed lines are the cross-family read-outs that the loss does not include.
Fig. 2: Configurations of the RWKV → Pythia and Pythia → RWKV chimera models.
Fig. 3: Correspondence over all 32×32 layer pairs of the unstandardized residuals h taken without the adapter. Left: linear CKA. Right: fitted R2 of per-layer linear maps. Dashed lines mark same-index layers, and the two panels use different color scales.
Fig. 4: Relation between the shared latent space and the four read-outs. Read-outs of the same color share a decoder.
ARC
SciQ
R → P
P → R
R → P
P → R
L
Ad
Af
Ad
Af
Ad
Af
Ad
Af
4
59.6
41.9
62.2
60.2
70.9
47.5
69.2
69.6
8
59.0
43.1
57.3
56.0
67.7
49.1
64.6
68.0
16
48.6
34.8
57.5
57.5
54.5
35.9
68.0
69.6
24
47.0
39.4
59.6
59.2
51.1
42.7
69.9
68.2
TABLE I: Length-normalized accuracy (%) of the adapter (Ad) and of the affine control for accuracy evaluation (Af). R → P: RWKV → Pythia; P → R: Pythia → RWKV.
ARC
SciQ
L
R → P
P → R
R → P
P → R
4
59.6→31.9
62.2→31.8
70.9→31.3
69.2→32.6
8
59.0→29.9
57.3→35.5
67.7→31.4
64.6→37.0
16
48.6→28.5
57.5→43.7
54.5→31.7
68.0→47.3
24
47.0→34.8
59.6→48.4
51.1→40.0
69.9→49.1
TABLE II: Change in length-normalized accuracy (%) when the nonlinear part is disabled, α:1→0 . R → P: RWKV → Pythia; P → R: Pythia → RWKV.
R self
R → P
P self
P → R
L
ARC
SciQ
ARC
SciQ
ARC
SciQ
ARC
SciQ
4
62.8
68.5
59.6
70.9
62.2
74.8
62.2
69.2
8
59.8
66.7
59.0
67.7
62.7
74.8
57.3
64.6
16
54.8
59.8
48.6
54.5
57.4
71.1
57.5
68.0
24
52.2
57.9
47.0
51.1
59.4
70.1
59.6
69.9
TABLE III: Length-normalized accuracy (%) of the four configurations. The maximum of each row and task is in bold. R self: RWKV self-reconstruction control; P self: Pythia self-reconstruction control; R → P: RWKV → Pythia; P → R: Pythia → RWKV.
ARC
SciQ
vs RWKV
Dir.
vs RWKV
Dir.
L
R → P
P → R
P → R − R → P
R → P
P → R
P → R − R → P
4
−2.2
+0.4
+2.6
+3.9
+2.2
−1.7
( 0.042 )
( 0.66 )
( 0.015 )
( 0.018 )
( 0.14 )
( 0.28 )
8
−2.8
−4.5
−1.7
+0.7
−2.4
−3.1
( 0.007 )
( <10−4 )
( 0.11 )
( 0.69 )
( 0.081 )
( 0.052 )
TABLE IV: Accuracy differences (points) between the main configurations and McNemar exact-test p -values in parentheses. R → P: RWKV → Pythia; P → R: Pythia → RWKV.
Config.
L
Trans- former layers
KV/token (KiB)
KV@4k (GiB)
Reduc. (%)
ARC/ SciQ
P alone
–
32
512
2.00
0
65.3/81.0
R → P
4
27
432
1.69
15.6
59.6/70.9
R → P
8
23
368
1.44
28.1
59.0/67.7
R → P
16
15
240
0.94
53.1
48.6/54.5
R → P
24
7
112
0.44
78.1
47.0/51.1
P → R
24
25
400
1.56
21.9
59.6/69.9
TABLE V: Accuracy and inference memory. KV@4k: KV cache size at context length 4096 . P alone: Pythia; R alone: RWKV; R → P: RWKV → Pythia; P → R: Pythia → RWKV.
Fig. 5: SciQ accuracy and KV cache reduction of RWKV → Pythia and Pythia → RWKV. The green line connects the configurations in the non-dominated set.
ARC
SciQ
Config.
L
Trans- former
Diff.
p
Diff.
p
P → R
4
5
+35.4
<10−4
+38.7
<10−4
R → P
24
7
+18.5
<10−4
+18.7
<10−4
P → R
8
9
+27.6
<10−4
+31.6
<10−4
R → P
16
15
+11.2
<10−4
+12.8
<10−4
P → R
16
17
+16.9
<10−4
+17.9
<10−4
TABLE VI: Differences in length-normalized accuracy (points, chimera model minus early exit) against a Pythia early exit that keeps the same number of Transformer layers, with McNemar exact-test p -values. Transformer: number of Transformer layers kept. R → P: RWKV → Pythia; P → R: Pythia → RWKV.
Fig. 6: Perplexity on Alpaca and WikiText at context length 512 . The vertical axis is logarithmic.
Prompt (verbatim)
Pythia alone
RWKV → Pythia ( L=4 )
RWKV alone
Pythia → RWKV ( L=4 )
### Instruction: What is the capital of France? ### Response:
“The capital of France is Paris.”
“The capital of France is Paris.” ✓
“Paris.”
“Paris.” ✓
The three primary colors are
“red, yellow, and blue. These colors are often used in combination to create a wide variety of colors.”
“red, yellow, and blue.” ✓
“red, blue, and yellow.” then drifts into a Q&A asking for green, the opposite color of red
“red, blue, and yellow.” then drifts into a Bob/Alice dialogue about green
Question: Water is made of hydrogen and Answer:
“Hydrogen is made of hydrogen and oxygen.”
“Hydrogen.” wrong
“hydrogen.” wrong
“hydrogen.” wrong
### Instruction: Explain what a black hole is in one sentence. ### Response:
“A black hole is a region of spacetime where gravity is so strong that nothing, not even light, can escape its pull.”
“A black hole is a region in space where gravity is so strong that nothing, not even light, can escape its pull.” ✓
“A black hole is a region of space where the gravitational pull is so strong that nothing, not even light, can escape.”
“A black hole is a type of celestial object that is formed when a star collapses and collapses into a black hole.” circular
TABLE VII: Generation examples of the two parent models and of RWKV → Pythia and Pythia → RWKV at L=4 .
We present a feedforward graph architecture in which heterogeneous frozen large language models serve as computational nodes, communicating through a shared continuous latent space via learned linear projections. Building on recent work demonstrating geometric compatibility between independently trained LLM latent spaces~\cite{armstrong2026thinking}, we extend this finding from static two-model steering to end-to-end trainable multi-node graphs, where projection matrices are optimized jointly via backpropagation through residual stream injection hooks. Three small frozen models (Llama-3.2-1B, Qwen2.5-1.5B, Gemma-2-2B) encode the input into a shared latent space whose aggregate signal is injected into two larger frozen models (Phi-3-mini, Mistral-7B), whose representations feed a lightweight cross-attention output node. With only 17.6M trainable parameters against approximately 12B frozen, the architecture achieves 87.3% on ARC-Challenge, 82.8% on OpenBookQA, and 67.2% on MMLU, outperforming the best single constituent model by 11.4, 6.2, and 1.2 percentage points respectively, and outperforming parameter-matched learned classifiers on frozen single models by 9.1, 5.2, and 6.7 points. Gradient flow through multiple frozen model boundaries is empirically verified to be tractable, and the output node develops selective routing behavior across layer-2 nodes without explicit supervision.
Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee
Department of Computer Science University of Houston Houston, TX 77204
Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model. We present a Universal Activation Bus, a framework that provides a common activation interface across compatible language models. Using a small set of source models, we learn a shared dense space together with one lightweight linear encoder--decoder adapter pair per model. After source training, the interface is frozen; a new model joins by fitting only its adapter pair on unlabeled matched text. The resulting interface allows activation-based tools to be shared across connected models, including common probes and SAE features as well as access to an NLA originally trained for a different model. Across five models, semantically related texts form consistent neighborhoods in the shared space, and an onboarded model reuses these tools effectively without retraining them. We further show that an intermediate activation from one model can be used by another model's frozen upper layers to produce predictions. These results establish a stable, model-wise activation contract for reusable tools across compatible language models.
Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han
Department of Artificial Intelligence, Yonsei University · Department of Computer Science, Yonsei University
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most 0.60 points, even when communication adds 15.44 points over the receiver alone. Thus the interface can supply the gain while making the sharer dispensable. Draft-KV instead sends the key-value states formed while the sharer drafts an answer to the current question. Linear projections place these states in a side memory read through a gated attention branch, and progressive training moves from message reconstruction to answer supervision under a guard on harm from mismatched messages. Both models remain frozen and the interface trains 1.05M parameters, 348x fewer than C2C. With a Qwen3-8B sharer, a frozen Qwen2.5-0.5B-Instruct receiver reaches 78.04% on MMLU-Redux, versus 37.45% alone and 36.40% with reassigned messages. At fixed interface size, scaling the sharer from 0.6B to 8B raises accuracy from 46.11% to 78.04%; communication also transfers to held-out tasks and can exceed both models when each holds different evidence.
Linquan Wu, Shichang Meng, Tianxiang Jiang +7
City University of Hong Kong · University of Science and Technology of China · University of Electronic Science and Technology of China +3