Organizations: School of Foreign Languages, Northwest University, Xi’an, China · Department of Quantitative Linguistics, University of Tübingen, Tübingen, Germany
How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.
recite a lesson for memory, endorse back, unlucky, recite, violate
Table 1 : Examples of relatively transparent (left) and opaque (right) compounds with character 书 ( shu1 , ‘book’).
Figure 2 : Core concepts of CAM. (a) A shift vector for a given constituent, here 水 ( shui3 , ‘water’), is obtained by subtracting the embedding of the other constituent, here 洪 ( hong2 , ‘huge’), from the compound embedding, here 洪水 ( hong2shui3 , ‘flood water’). (b) Compounds with 水 (‘water’, in blue) cluster (using t-SNE) in the embedding space. The points in grey represent compounds with nine other high-frequency monomorphemic words: 火 (‘fire’), 金 (‘gold’), 木 (‘wood’), 土 (‘earth’), 天 (‘sky’), 道 (‘road’), 心 (‘heart’), 人 (‘man’), and 名 (‘name’). (c) The corresponding shift vectors of 水 likewise cluster in the corresponding shift space.
Figure 5 : Geometric illustration of CAM. Constituent embeddings sc1 and sc2 act as anchor points, while centroid shift vectors dˉc1,1 and dˉc2,2 provide position-specific adjustments. The mean of the two transcended contituents ((sc1+dˉc1,1)+(sc2+dˉc2,2))/2 yields the predicted compound representation s^w .
Figure 6 : Prediction accuracy for CAM and CAOSS models for training and test data. For each of 30 runs, data were randomly split into training (90%) and test (10%) sets.
Figure 7 : Prediction accuracy of CAM and CAOSS under a frequency-based split with the lowest-frequency 10% of compounds in the test set and the remaining 90% in the training set.
Figure 8 : Boxplots of accuracy for CAM (red) and CAOSS (blue) models trained on only the words with a given length (2, 3, and 4 characters). CAM outperformed CAOSS except for the held-out data with three-character compounds.
Figure 9 : Prediction accuracy of CAM and CAOSS for compounds of different word lengths under a frequency-based split, with the lowest-frequency 10% of compounds assigned to the test set and the remaining 90% to the training set. CAM significantly outperformed CAOSS only for the held-out data of two-character compounds.
Word length
Dataset
CAM
CAOSS
χ2(1)
p
2-char
Train
0.37
0.13
4455.7
<.0001
Test
0.17
0.09
111.26
<.0001
3-char
Train
0.94
0.71
1007.8
<.0001
Test
0.37
0.41
1.40
0.24
4-char
Train
0.99
0.94
52.90
<.0001
Test
0.69
0.56
2.69
0.10
Table 2 : Comparison of CAM and CAOSS accuracy for training and test splits based on frequency, using proportions tests.
Figure 10 : Boxplots for CAM (red) and CAOSS (blue) accuracies nouns by word length, for 30 cross-validation runs.
Figure 11 : Distributions of the vector lengths (L2-norms) for CAOSS (left panel) and CAM (right panel), for transcended (blue) and observed (red) constituent embeddings, as well as for the predicted and observed compound vectors.
Figure 12 : Absolute-value heatmap of the global CAOSS transformation matrix W for two-character Mandarin compounds. The two blocks correspond to the constituent-specific sub-matrices W1 and W2 . Both matrices are diagonally dominant, with brighter diagonal entries indicating larger transformation weights. The diagonal elements account for approximately 72.8% of the Frobenius norm of W .
A. parametric coefficients
Estimate
Std. Error
t-value
p-value
Intercept
-1.1593
0.0333
-34.7889
< 0.0001
Pc1CAOSS
-0.0727
0.0226
-3.2211
0.0013
Pc2CAOSS
-0.1870
0.0312
-5.9939
< 0.0001
B. smooth terms
edf
Ref.df
F-value
p-value
s(logf(w))
5.3077
6.3023
933.7357
< 0.0001
te ( V(c2),SCCAOSS )
13.1744
15.6219
5.1673
< 0.0001
Table 4 : Summary of a GAM fitted to visual lexical decision latencies with predictors derived from the CAOSS model (AIC: -11570.1).
Figure 13 : Partial effects of the smooth terms in the GAM fitted to visual lexical decision latencies using predictors derived from the CAOSS model. The left panel shows the nonlinear effect of log-transformed word frequency, with the effect levelling off at higher frequency levels. The middle panel displays the observed distribution of the two predictors. The right panel shows the contour plot of the tensor smooth for the interaction between right constituent family size and cosine similarity. Colours indicate predicted response latencies, with darker blue regions representing shorter latencies and red regions representing longer latencies.
df
AIC
Δ AIC
f(w)
30.4
-7075.4
4494.7
V(c2)
19.4
-11512.7
57.4
Pc2CAOSS
23.3
-11537.2
32.9
Pc1CAOSS,Pc2CAOSS
22.8
-11537.3
32.8
SCCAOSS
18.5
-11560.3
9.8
Pc1CAOSS
27.6
-11561.2
8.9
Table 5 : Variable importance, estimated by the change in AIC when a predictor is removed from the GAM model that is informed by CAOSS-based predictors. ( 22 ).
Figure 14 : The geometry of CAOSS proximity measures. Left panel: transcended vectors of equal length; right panel: transcended vectors of different lengths.
A. parametric coefficients
Estimate
Std. Error
t-value
p-value
Intercept
-1.3555
0.0013
-1018.1199
< 0.0001
B. smooth terms
edf
Ref.df
F-value
p-value
s(log(f(w))
5.4745
6.4724
852.9476
< 0.0001
te(V(c2),SCCAM
14.5436
17.2955
3.4391
< 0.0001
te(Pc1CAM,Prc2CAM)
3.8423
4.5405
17.2632
< 0.0001
Table 6 : Summary of a GAM fitted to visual lexical decision latencies with predictors grounded in CAM (AIC: -11597.45).
Figure 15 : Partial effects of the smooth terms in a GAM fitted to visual lexical decision latencies using predictors derived from the CAM model. The first panel in first row shows the nonlinear effect of log-transformed word frequency. The middle and right panels in the first row display the observed distributions of the predictors involved in the two tensor-product smooths: right constituent family size ( V(c2) ) with cosine similarity ( SCCAM ), and the proximity measures of the first and second constituents ( Pc1CAM and Pc2CAM ), respectively. The second row presents the corresponding contour plots of the tensor smooths. Colours indicate predicted response latencies, with bluer shades representing shorter response latencies and lighter/redder shades representing longer response latencies.
df
AIC
Δ AIC
f(w)
25.6
-7333.5
4264.0
Pc1CAM,Pc2CAM
22.7
-11527.9
69.6
Pc1CAM
23.6
-11538.9
58.6
V(c2)
14.1
-11567.1
30.4
SCCAM
19.4
-11584.9
12.5
Pc2CAM
29.1
-11589.8
7.7
Table 7 : Variable importance, estimated by the change in AIC when a predictor is removed from the GAM model based on CAM predictors ( 24 ).
Figure 16 : Comparison of variable importances for CAM-based (blue) and CAOSS-based (orange) predictors, as well as second constituent family size, assessed with the increase in AIC when a predictor is withheld from the GAM.