Platonic Task Arithmetic
Organizations: TelePIX
Abstract
Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows of one shared, model-agnostic object, which we call the platonic task vector. To make it operational for models that pair an image or audio encoder with a text encoder, we introduce Universal Task Descriptors: matrices whose shape is independent of architecture and embedding dimension, which record a task's functional effect and support addition and negation as matrix operations. Transferring a descriptor into a target means editing the target until it reproduces the descriptor on the task's unlabeled probe images and class-name prompts, requiring no per-image labels. We realize this edit in two ways. First, the descriptor factorizes into a shift field on image embeddings, so a single least-squares solve yields a linear operator that folds into the target's last layer as a weight edit; by linearity, a bank of such operators admits any composition at any strength as a signed sum. Second, a low-rank adapter trained on the same objective reaches every layer and fits compositions jointly, at the cost of one optimization per edit. Heterogeneous models share this object only partially, with a model-specific residual comparable in norm to the shared component, yet cross-model transfer still retains 74-80 percent of the gain of the target's own descriptors. Experiments across six model families, eight classification tasks, and an audio-text setting show that task knowledge transfers and composes across heterogeneous models under both realizations.
Figures & tables
| Method | Task A | Task B | CIFAR-100 | ImageNet |
|---|---|---|---|---|
| Source and target are the same model ( cells) | ||||
| Zero-shot | ||||
| Weight-space arithmetic | ||||
| Ours, closed form | ||||
| Ours, adapter | ||||
| Different models within one architecture ( cells) | ||||
| Method | Removed task | CIFAR-100 | ImageNet |
|---|---|---|---|
| Source and target are the same model ( cells) | |||
| Zero-shot | |||
| Weight-space negation | |||
| Ours, closed form | |||
| Ours, adapter | |||
| Different models within one architecture ( cells) | |||
| Assignment | Route | RESISC45 | UCMerced | AID | Mean |
|---|---|---|---|---|---|
| None | Zero-shot | ||||
| Closed form | |||||
| Adapter | |||||
| Closed form | |||||
| Adapter | |||||
| alone | Closed form |
| Method | Removed | Retained | UrbanSound8K |
|---|---|---|---|
| Zero-shot | |||
| Source and target are the same model ( cells) | |||
| Ours, closed form | |||
| Ours, adapter | |||
| Source and target are different models ( cells) | |||
| Ours, closed form | |||
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Description |
|---|---|
| Models, encoders and parameters | |
| The -th vision–language model, . | |
| Target model receiving the edit. | |
| , | Task-term index and the source assignment . |
| Vision and text encoders of , with frozen. | |
| Pre-trained and fine-tuned vision encoders of . | |
| Target | Cosine | Relative difference, double | Relative difference, single |
|---|---|---|---|
| CLIP-L | |||
| OpenCLIP-L | |||
| MetaCLIP-L | |||
| EVA-CLIP-L |
| Target | Guaranteed | Changed | Bound | Observed | Single |
|---|---|---|---|---|---|
| CLIP-L | 48.1 | 9.5 | 23.78 | 1.30 | 71.83 |
| OpenCLIP-L | 49.5 | 9.0 | 16.82 | -0.29 | 66.30 |
| MetaCLIP-L | 46.5 | 9.1 | 21.52 | 1.21 | 68.03 |
| EVA-CLIP-L | 47.2 | 11.8 | 28.60 | 3.58 | 75.85 |
| Pooled | 47.8 | 9.9 | 22.68 | 1.45 | 70.50 |
| Agreement measured on | Mean agreement | Rank correlation | Best source identified |
|---|---|---|---|
| Descriptors | of | ||
| Operators | of | ||
| Shift fields on the probes | of |
| EuroSAT | DTD | Cars | GTSRB | MNIST | RESISC45 | SVHN | SUN397 | |
|---|---|---|---|---|---|---|---|---|
| Descriptor agreement | ||||||||
| Operator agreement | ||||||||
| Shift-field agreement | ||||||||
| Gain forfeited |
| Rank cap | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | Prompts | Selected | ||||||||
| EuroSAT | — | — | — | — | — | |||||
| DTD | — | — | — | |||||||
| Cars | — | |||||||||
| GTSRB | — | — | — | |||||||
| MNIST | — | — | — | — | — | |||||
| Family | Identifier | Adapted modules | |
| CLIP | openai/clip-vit-large-patch14 | q_proj , v_proj | |
| OpenCLIP | laion/CLIP-ViT-L-14-laion2B-s32B-b82K | q_proj , v_proj | |
| MetaCLIP | facebook/metaclip-l14-400m | q_proj , v_proj | |
| EVA-CLIP | EVA02-L-14 (via open_clip ) | attn.{q,v}_proj | |
| SigLIP | google/siglip-large-patch16-256 | q_proj , v_proj | |
| SigLIP2 | google/siglip2-large-patch16-256 | q_proj , v_proj |
| Method | Task A | Task B | CIFAR-100 | ImageNet | |
|---|---|---|---|---|---|
| Ours | |||||
| Single-A | |||||
| Single-B | |||||
| Ours | |||||
| Single-A | |||||
| Single-B |
| Target | Composed tasks | CIFAR-100 | ImageNet | Zero-shot CIFAR-100 |
|---|---|---|---|---|
| CLIP-L | ||||
| OpenCLIP-L | ||||
| MetaCLIP-L | ||||
| EVA-CLIP-L | ||||
| SigLIP-L | ||||
| SigLIP2-L |
| Task | Ours | Oracle | Fraction retained |
|---|---|---|---|
| EuroSAT | |||
| DTD | |||
| Cars | — | ||
| GTSRB | |||
| MNIST | |||
| RESISC45 |
| Task | CLIP-L | OpenCLIP-L | MetaCLIP-L | EVA-CLIP-L | Mean |
|---|---|---|---|---|---|
| CIFAR-100 | |||||
| EuroSAT | |||||
| DTD | |||||
| Cars | |||||
| GTSRB | |||||
| MNIST | |||||
| Cross-model source | Same-model source | ||||||
|---|---|---|---|---|---|---|---|
| Removed task | CIFAR-100 | ImageNet | Removed task | CIFAR-100 | ImageNet | ||
| Addition at | EuroSAT | DTD | CIFAR-100 | ImageNet |
|---|---|---|---|---|
| Zero-shot | ||||
| Two sources, one per task ( cells) | ||||
| the study’s own sources | ||||
| One source, both tasks ( cells) | ||||
| Target’s own sources ( cells) | ||||
| the study’s own sources |
| Assignment | Method | RESISC45 | UCMerced | AID | Mean | CIFAR-100 |
|---|---|---|---|---|---|---|
| None | Zero-shot | |||||
| Within-model | Composition | |||||
| Within-model | Single | |||||
| Within-model | Oracle (target’s own source) | |||||
| Cross-model, CLIP | Composition | |||||
| Cross-model, MetaCLIP | Composition |
| Arbitrary halves, photographs and sketches | Semantic split, photographs and paintings | |||
| Method | Held-out task | CIFAR-100 | Held-out task | CIFAR-100 |
| Zero-shot | ||||
| Weight-space task arithmetic | ||||
| First term alone | ||||
| Third term alone | ||||
| Analogy, | ||||
| Method | Structural requirement | Target-side signal | Operations | Different families |
| Weight-space task arithmetic [ Ilharco et al., 2022 ] | Identical parameter space | None | Addition, negation, analogy | n/a |
| GradFix [ Rinaldi et al., 2026a ] | Matching parameter coordinates | Labeled examples and their gradients | Addition, negation | Partly |
| Theseus [ Rinaldi et al., 2026b ] | Corresponding submodules, width mismatch allowed | Calibration activations | Addition | Partly |
| Ours, closed form | None beyond a cross-modal affinity and an affine last layer writing into the embedding | Unlabeled probe images and prompts | Addition, negation | Yes |
| Ours, adapter | None beyond a cross-modal affinity | Unlabeled probe images and prompts, one optimization per edit | Addition, negation | Yes |
| Gradient ascent [ Thudi et al., 2022 ] | None, since it trains the target directly | Labels for the removed task | Negation | n/a |
| Method | Strength | Removed task | CIFAR-100 | ImageNet |
|---|---|---|---|---|
| Against weight-space negation ( cells) | ||||
| Weight-space negation | ||||
| Weight-space negation | ||||
| Weight-space negation | ||||
| Ours, closed form, same model | ||||
| Ours, closed form, same model | ||||
| Target | EuroSAT | DTD | Cars | GTSRB | MNIST | RESISC45 | SVHN | SUN397 |
|---|---|---|---|---|---|---|---|---|
| CLIP-L | ||||||||
| OpenCLIP-L | ||||||||
| MetaCLIP-L | ||||||||
| EVA-CLIP-L |
| Addition | EuroSAT | DTD | CIFAR-100 | ImageNet |
|---|---|---|---|---|
| Zero-shot | ||||
| Ours, closed form | ||||
| Ours, adapter | ||||
| Oracle, closed form | ||||
| Affinity distillation | ||||
| Posterior distillation |
| Target | Method | Task A | Task B | CIFAR-100 | ImageNet |
|---|---|---|---|---|---|
| SigLIP-L | Zero-shot | ||||
| Random-UTD | – | ||||
| Head edit | |||||
| Embedding operator | |||||
| Oracle | |||||
| SigLIP2-L | Zero-shot |
| Addition at | Negation at | Negation at | ||||||
|---|---|---|---|---|---|---|---|---|
| Target | Composed | CIFAR-100 | ImageNet | Removed | CIFAR-100 | Removed | CIFAR-100 | |
| CLIP-B | ||||||||
| MetaCLIP-B | ||||||||
| OpenCLIP-H | ||||||||
| MetaCLIP-H | ||||||||
| Large targets, reference | ||||||||
| Method | ESC-50 | GTZAN | UrbanSound8K |
|---|---|---|---|
| Zero-shot | |||
| Random-UTD | |||
| Ours, closed form | |||
| Oracle | |||
| Zero-shot, adapter runs | |||
| Ours, adapter |
| Source (ESC-50) | Source (GTZAN) | Target | Ours |
|---|---|---|---|
| unfused | large | large-ms | / / |
| unfused | large-ms | large | / / |
| large | large-ms | unfused | / / |
| large | unfused | large-ms | / / |
| large-ms | unfused | large | / / |
| large-ms | large | unfused | / / |