Weight compression can alter accelerator placement as well as memory traffic, complicating the interpretation of inference speedups. We investigate this interaction on the Apple Neural Engine through the public Core ML deployment path. Five independently trained language-model checkpoints span two architectures and dense fp16, int8, and ternary weights encoded with two-bit lookup tables. We combine compiler device plans, synchronized memory-controller measurements, and compute-unit exclusion controls for a fixed single-token forward workload. On an M1, the smaller fp16 export executes on the CPU despite permitting ANE execution, whereas its compressed counterparts exhibit ANE activity. The int8 export reduces warm forward latency by a factor of 1.9. The larger fp16 export also uses the ANE, indicating that dense encoding alone does not determine placement. Separate M3 energy measurements support the same direction of change. These results establish encoding-dependent placement in the measured deployment stack and show that compression comparisons require joint measurement of backend selection, latency, and traffic.
Figures & tables
Comparison
Held fixed
Changed
Recorded
Weight encoding
Architecture within each family; input
Training precision and exported weights
Device plan, latency, bytes
Device exclusion
Compiled model and input
CPU only, CPU+ANE, CPU+GPU
Latency and counter signal
Normalization dtype
Arithmetic; fp16 input/output
fp16 or fp32 intermediates
Eligibility and output error
Table 1: Experimental comparisons. Encoding arms use separately trained weights; device controls reuse each compiled model.
Model
Hidden / layers
Total M
Quantized M
Accounted MB
q25-fp16
512 / 10
27.943
0.000
55.886
q25-int8
512 / 10
27.943
25.821
30.065
q25-ternary
512 / 10
27.943
25.821
10.699
q50-fp16
640 / 12
51.155
0.000
102.311
q50-ternary
640 / 12
51.155
48.497
17.442
Table 2: Principal checkpoints and mixed-width inventory. A zero quantized count denotes the dense fp16 arm.
Model
ANE-pref. / assigned
CPU ms
CPU+ANE ms
CPU+GPU ms
q25-fp16
0/331
1.454
1.458
3.256
q25-int8
324/335
1.419
0.770
8.631
q25-ternary
324/335
1.508
0.638
16.255
q50-fp16
380/387
2.787
2.079
5.071
q50-ternary
380/391
2.757
0.909
19.667
Table 3: M1 measurements. Entries are medians of per-window median warm prediction times; requested units permit fallback.
M3 cell
ANE energy mJ
Prediction ms
q25-fp16
0
1.125
q25-int8
2200
0.645
int8 / CPU+ANE
2148
0.640
int8 / CPU only
0
1.130
int8 / CPU+GPU
0
3.760
Table 4: M3 placement and compute-unit controls. Energy Model/ANE readings in millijoules indicate accelerator activity; they are not energy per prediction.
Model
MB / call
Window range MB
Ratio to inventory
GB/s
q25-fp16
0.000
0.000–0.000
0.000
0.00
q25-int8
28.463
28.461–28.476
0.947
36.25
q25-ternary
11.134
11.132–11.144
1.041
17.31
q50-fp16
102.795
102.790–102.799
1.005
48.00
q50-ternary
17.931
17.927–17.936
1.028
19.62
Table 5: M1 ANE memory traffic (read + write bytes) for CPU_AND_NE. Ratios use the accounted weight bytes in the checkpoint table; the zero cell is a placement null.
Graph
ANE eligible / assigned
Max abs. error
Max rel. L2
Checks
rmsnorm fp16
7/7
0.002362
0.000530
20
rmsnorm fp32
0/8
0.000975
0.000283
20
add fp16
2/2
0.000000
0.000000
20
add fp32
0/3
0.000000
0.000000
20
Table 6: Matched dtype controls. Error maxima are against the numerical reference across the retained inputs and requested devices.
Export
Package MB
Best CPU+ANE ms
Best CPU ms
q50-fp16
102.5
1.903
2.543
q50-int8
51.5
1.105
2.502
q25-fp16
56.0
1.346
1.349
q50-ternary
17.6
0.865
2.524
q25-int8
28.2
0.765
1.350
a25-dense-fp16
49.0
1.103
1.119
Table 7: Earlier sequence-one export timings. Package size differs from the accounted weight inventory; replication is weaker than in the main experiment.