Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.
Figures & tables
Figure 1 : Motivation for clipping calibration in analog IMC. Clipping changes both operand quantization error and the output-referred ADC quantization error, motivating their joint optimization.
Figure 2 : Output-error analysis for LLaMA-3.2-3B across projections. Each point reports the median normalized MSE across 28 decoder layers, and shaded regions denote the 25th–75th percentiles. The first three panels isolate activation-, weight-, and ADC-induced output errors, while the final panel measures the total IMC output error directly. Lower is better.
Figure 3 : Calibration and optimization efficiency on LLaMA-3.2-3B. Left: IMC-CLINIC achieves a better WikiText-2 PPL–calibration-time trade-off than direct ADC-aware alternating search. Right: Its safeguarded Newton optimizer reaches 95% of the post-initialization loss improvement faster than projected gradient descent ( 72 ms vs. 1.10 s median).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Empirical validation of cross-coordinate error decorrelation on LLaMA-3.2-3B. Left: Pairwise Pearson correlations between centered MAC errors for a representative output channel in the layer-13 q projection. Right: Ratio of the decorrelated approximation to the exact operand MSE across decoder layers. Faint points denote individual layers, diamonds denote the median, shaded regions indicate the 25th–75th percentiles, and the dashed line denotes exact agreement.
Figure 5 : Empirical validation of the two-term second-moment approximation on LLaMA-3.2-3B. Left: Signed contributions of the six exact terms for the representative layer-13 q projection, where the two retained terms account for 95.5% of the total. Right: Ratio of the retained two-term sum to the exact six-term sum across decoder layers. Faint points denote individual layers, diamonds denote the median, shaded regions indicate the 25th–75th percentiles, and the dashed line denotes exact agreement.
Figure 6 : Empirical validation of the signed-bias term on LLaMA-3.2-3B. Left: Mean signed MAC error E[δyi] across MAC coordinates for a representative output channel in the layer-13 q projection. Right: Ratio of predicted to measured operand-output MSE across decoder layers, comparing the surrogate without and with Lbias . Faint points denote individual layers, diamonds denote the median, shaded regions indicate the 25th–75th percentiles, and the dashed line denotes exact agreement.
Figure 7 : Empirical validation of operand–ADC error decorrelation on LLaMA-3.2-3B. Left: Operand-induced output error versus the additional ADC error for the representative layer-13 q projection, with Pearson correlation r=−0.001 . Right: Pearson correlation across all 28 decoder layers and projection types. Faint points denote individual layers, diamonds denote the median, shaded regions indicate the 25th–75th percentiles, and the dashed line denotes zero correlation.
Figure 8 : PSD regions of the full surrogate Hessian for the q , o , up , and down projections in decoder block 13 of LLaMA-3.2-3B. Blue points are sampled clipping-factor configurations with λmin>0 , gray points have λmin<0 , and the red surface marks the λmin=0 boundary. The insets show the optimization trajectories (black dashed lines) from initialization (blue circles) to the calibrated solutions (red stars); both endpoints lie in the PSD region in all four projections.
Property
LLaMA-3.2-3B
Qwen3-4B
Single connected PSD component
196/196 projs
252/252 projs
Initialization in PSD component
196/196 projs
252/252 projs
Calibrated solution in PSD component
196/196 projs
252/252 projs
Accepted trajectory remains in PSD component
196/196 projs
252/252 projs
Appendix
Table 2 : Full-Hessian PSD-region validation across all projections of LLaMA-3.2-3B and Qwen3-4B. Entries report the number of projections satisfying each property.
Figure 9 : Relative mismatch between the analytical surrogate and empirically measured MatMul output MSE for LLaMA-3.2-3B at the calibrated solutions. Points denote individual layers, solid lines show the median across all 28 layers, and shaded regions indicate the 25th–75th percentiles. Each panel reports the indicated error component across projection types.
Figure 10 : Output-error analysis for the three additional models. Each subfigure shows normalized MSE from activation quantization, weight quantization, and ADC quantization, followed by the total IMC output error across projection types. Points report the median across decoder layers, and shaded bands show the interquartile range. Lower is better.
Figure 11 : Clipping factors for LLaMA-3.2-3B under the default 9-bit ADC configuration. Results are shown for the upper activation factor γ , lower activation factor β , and weight factor α across projection types. Markers denote the mean across decoder layers, and shaded regions indicate ±1 standard deviation. A clipping factor of 1 corresponds to no clipping.
Model
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
Calib. Time ( ↓ )
LLaMA-3.2-3B
No clipping
175.70
0.515
0.268
0.533
0.241
0.527
0.293
0.310
0.384
–
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
32.9 min
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
3.3 min
W8A8 no-IMC
7.82
0.698
0.402
0.778
0.464
0.745
0.723
0.740
0.650
–
FP16 baseline
7.81
0.694
0.408
0.781
0.462
0.742
0.721
0.741
0.650
–
LLaMA-3.1-8B
No clipping
127.28
0.460
0.274
0.572
0.263
0.489
0.363
0.357
0.397
–
Appendix
Table 3 : Full WikiText-2 perplexity (PPL), seven-task zero-shot accuracy, and calibration time across models and inference configurations. Avg. is the unweighted mean task accuracy; a dash indicates that calibration time is unavailable or inapplicable.
ADC bits
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
8
No clipping
29717.68
0.502
0.294
0.505
0.271
0.412
0.256
0.262
0.358
W/A grid search
1572.10
0.510
0.262
0.504
0.268
0.395
0.259
0.263
0.352
IMC-CLINIC
42.63
0.530
0.244
0.585
0.242
0.531
0.402
0.376
0.416
9 ∗
No clipping
175.70
0.515
0.268
0.533
0.241
0.527
0.293
0.310
0.384
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
Appendix
Table 4 : LLaMA-3.2-3B results across ADC precision.
IMC rows
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
128
No clipping
10.36
0.627
0.404
0.737
0.419
0.647
0.660
0.696
0.599
W/A grid search
9.39
0.669
0.400
0.751
0.406
0.657
0.677
0.699
0.608
IMC-CLINIC
9.04
0.680
0.408
0.749
0.434
0.706
0.692
0.712
0.626
256
No clipping
15.86
0.564
0.326
0.681
0.331
0.518
0.528
0.593
0.506
W/A grid search
11.42
0.647
0.358
0.742
0.398
0.607
0.657
0.661
0.581
IMC-CLINIC
10.03
0.648
0.384
0.744
0.401
0.654
0.660
0.688
0.597
Appendix
Table 5 : LLaMA-3.2-3B results across IMC row counts.
IMC columns
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
8
No clipping
175.32
0.505
0.228
0.542
0.223
0.538
0.299
0.311
0.378
W/A grid search
19.12
0.565
0.296
0.672
0.317
0.515
0.529
0.525
0.488
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
16
No clipping
175.32
0.504
0.228
0.542
0.223
0.538
0.299
0.311
0.378
W/A grid search
19.12
0.565
0.296
0.672
0.317
0.515
0.529
0.525
0.488
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
Appendix
Table 6 : LLaMA-3.2-3B results across IMC column counts.
Noise (LSB)
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
0 ∗
No clipping
175.70
0.515
0.268
0.533
0.241
0.527
0.293
0.310
0.384
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
0.1
No clipping
288.45
0.514
0.256
0.503
0.242
0.478
0.294
0.287
0.368
W/A grid search
21.70
0.552
0.300
0.599
0.299
0.536
0.485
0.490
0.466
IMC-CLINIC
12.58
0.614
0.352
0.679
0.345
0.643
0.601
0.636
0.553
Appendix
Table 7 : LLaMA-3.2-3B robustness to analog noise.
Samples
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
Calib. Time ( ↓ )
8 ∗
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
32.9 min
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
3.4 min
16
W/A grid search
N/E
N/E
N/E
N/E
N/E
N/E
N/E
N/E
N/E
>60 min (T/O)
IMC-CLINIC
12.21
0.631
0.350
0.716
0.373
0.623
0.642
0.654
0.570
3.9 min
32
W/A grid search
N/E
N/E
N/E
N/E
N/E
N/E
N/E
N/E
N/E
>60 min (T/O)
IMC-CLINIC
12.31
0.631
0.370
0.723
0.384
0.650
0.654
0.641
0.579
5.0 min
Appendix
Table 8 : LLaMA-3.2-3B results across calibration-set sizes. N/E indicates metrics not evaluated because calibration timed out; T/O marks a run exceeding the 60-minute limit.
Seq. length
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
Calib. Time ( ↓ )
512
W/A grid search
17.72
0.548
0.302
0.671
0.317
0.523
0.529
0.539
0.490
10.9 min
IMC-CLINIC
12.26
0.611
0.368
0.713
0.376
0.630
0.648
0.651
0.571
3.0 min
1024
W/A grid search
18.46
0.586
0.310
0.655
0.300
0.497
0.530
0.539
0.488
18.4 min
IMC-CLINIC
12.20
0.638
0.372
0.731
0.386
0.660
0.653
0.645
0.584
3.1 min
2048 ∗
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
32.9 min
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
3.4 min
Appendix
Table 9 : LLaMA-3.2-3B results across calibration sequence lengths.
Domain
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
WikiText-2 ∗
W/A grid search
19.12
0.566
0.294
0.672
0.317
0.515
0.528
0.525
0.488
IMC-CLINIC
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
C4
W/A grid search
17.73
0.568
0.316
0.676
0.306
0.542
0.546
0.559
0.502
IMC-CLINIC
12.32
0.627
0.376
0.720
0.381
0.642
0.643
0.646
0.576
Pile
W/A grid search
17.88
0.566
0.316
0.655
0.331
0.553
0.555
0.552
0.504
IMC-CLINIC
12.31
0.623
0.362
0.729
0.373
0.685
0.647
0.652
0.581
Appendix
Table 10 : LLaMA-3.2-3B results across calibration domains.
Figure 12 : Optimization convergence of IMC-CLINIC on LLaMA-3.2-3B. Left: Surrogate loss versus accepted Newton iteration for decoder block 13, normalized by the loss at the selected initialization, L(t)/L(0) . Right: Number of accepted Newton iterations at optimizer return across all 28 decoder blocks. Faint points denote individual decoder blocks, diamonds denote the median, and vertical ranges indicate the 25th–75th percentiles.
Method
Joint W/A
Signed bias
ADC term
PPL ( ↓ )
IMC-CLINIC
✓
✓
✓
12.27
w/o ADC term
✓
✓
127.91
w/o signed-bias term
✓
✓
12.76
Independent W/A
✓
✓
15.80
Appendix
Table 11 : Method ablations on LLaMA-3.2-3B under the 9-bit ADC setting. All variants use the same default initialization, and ADC quantization remains enabled during inference.
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
IMC-CLINIC + per-channel α
12.02
0.617
0.370
0.721
0.393
0.668
0.649
0.650
0.581
IMC-CLINIC (shared scalar α )
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
Appendix
Table 12 : LLaMA-3.2-3B results across weight-clipping granularities.
Method
PPL ( ↓ )
Wino
OBQA
PIQA
ARC-C
BoolQ
ARC-E
Hella
Avg. ( ↑ )
IMC-CLINIC +EF-weighted loss
11.88
0.619
0.364
0.716
0.370
0.622
0.631
0.657
0.568
IMC-CLINIC (unweighted loss)
12.27
0.622
0.368
0.722
0.367
0.665
0.630
0.651
0.575
Appendix
Table 13 : LLaMA-3.2-3B results with empirical-Fisher weighting for input-divergent projections.
Figure 13 : Branch-and-bound procedure for validating global 1% -optimality of the calibrated clipping factors under the surrogate objective.
Model
Q
K
V
O
Gate
Up
Down
Total
Median Time
LLaMA-3.2-3B
28/28
28/28
28/28
28/28
28/28
28/28
28/28
196/196
16.88 min
Qwen3-4B
36/36
36/36
36/36
36/36
36/36
36/36
36/36
252/252
18.58 min
Appendix
Table 14 : Optimality-validation results. Each entry under a projection type reports the number of projection-level clipping solutions certified to be within 1% of the global minimum of the calibration objective. Validation time is the median wall-clock time per projection and is incurred only for the optimality check, not during clipping calibration.
Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit quantized models. Existing CIM-oriented quantization methods mainly minimize ideal quantization error, ignoring the hardware noise floor and thus causing inefficient precision allocation. We propose NANQ, a noise-aware mixed-precision non-uniform quantization framework for analog CIM. NANQ models magnitude-dependent weight noise from measured responses of an eFlash CIM array and converts the noise profile into an adaptive quantization density, assigning finer resolution to low-noise regions while avoiding ineffective precision in noise-dominated regions. It further assigns layer-wise bit-widths by identifying each layer's precision saturation point under hardware noise using a unified threshold. On-chip experiments on an eFlash CIM SoC show that, under 2-bit weight-magnitude quantization, NANQ improves vision-model accuracy by 8.05 percentage points and reduces language-model PPL by 54.7% on average over PowerQuant. Mixed-precision NANQ captures most of the gains obtainable from additional quantization resources with only 3.2-3.8 equivalent bits.
Yizhe Chen, Wenshuai Yao, Saiya Wang +6
School of Integrated Circuit Science and Engineering, Beihang University, Beijing, China · School of Integrated Circuits, Peking University, Beijing, China · Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong SAR, China
Analog in-memory computing (AIMC) speeds up neural-network inference by doing the arithmetic directly inside a memory array, instead of shuttling weights back and forth between memory and a processor. This saves energy, but the physical devices that store the weights are imperfect: programming errors, electrical noise, limited-resolution converters, and outright broken cells all distort the computation, and every physical chip is distorted in its own way. A designer with several such chips available faces an uncomfortable choice: run all of them and combine the answers (safe, but wasteful of energy), or trust a single chip blindly (cheap, but with no guarantee on how often it is wrong). This paper introduces RACE-AIMC (Risk-Aware Certified Ensemble for AIMC), a framework that resolves this choice with statistics rather than guesswork. Offline, RACE-AIMC studies a pool of physical accelerators, picks the single best one for a given energy budget, and computes a mathematically exact upper bound on how often that accelerator will be wrong when it chooses to answer. Online, only that one accelerator is switched on; a lightweight check decides whether to accept its answer or defer to a fallback. In our simulations using a noisy weight mapping and multiple independent test runs, every certified bound stayed under a 10% error target (mean bound 7.83% +- 0.89%, with 70.88% +- 0.98% of inputs answered directly). The resulting system matches the accuracy of a clean digital baseline while cutting modeled energy use by 69.02% relative to always running every accelerator in the pool.
Analog in-memory computing (AIMC) performs computation directly within resistive crossbar arrays, offering an energy-efficient platform to scale large vision and language models. However, non-ideal analog device properties make the training on AIMC devices challenging. In particular, its update asymmetry can induce a systematic drift of weight updates towards a device-specific symmetric point (SP), which typically does not align with the optimum of the training objective. To mitigate this bias, most existing works assume the SP is known and pre-calibrate it to zero before training by setting the reference point as the SP. Nevertheless, calibrating AIMC devices requires costly pulse updates, and residual calibration error can directly degrade training performance. In this work, we present the first theoretical characterization of the pulse complexity of SP calibration and the resulting estimation error. We further propose a dynamic SP estimation method that tracks the SP during model training, and establishes its convergence guarantees. In addition, we develop an enhanced variant based on chopping and filtering techniques from digital signal processing. Numerical experiments demonstrate both the efficiency and effectiveness of the proposed method.
Quan Xiao, Jindan Li, Zhaoxian Wu +2
Department of Electrical and Computer Engineering, Cornell University, New York, NY · IBM T. J. Watson Research Center, Yorktown Heights, NY · Rensselaer Polytechnic Institute, Troy, NY