Radar hardware faults threaten automated perception, motivating accurate, compact diagnosis and understandable maintenance guidance. We introduce SCORE-LM, which couples a small scatterer-conditioned operator-response encoder (SCORE) to an adapted local language model. SCORE combines self-referenced complex trajectories, physical descriptors, and a selective state-space branch, with source-only self-supervision and directional fault inference. On eight capture-excluded Rad-R fault recordings, it achieves state-of-the-art performance within the evaluated nine-model comparison: 88.39% mean capture recall and 88.20% four-fault macro-F1 at ten frames. Its 39,520 radar inference coefficients are 119.7 times fewer than RadrNet-DS-CI's, while recall is 15.56 percentage points higher than this strongest competitor. In a separate low-label protocol, SCORE reaches 71.58% recall with one labeled source window per class. A nonlinear projector converts four frozen fault similarities into five soft tokens, linking compact diagnosis to class-conditioned maintenance guidance. On 75 development questions covering 24 radar windows, language adaptation raises correct-fault answers from 45 to 62 (60.0% to 82.7%) relative to removing the co-trained adapters, while retaining the same projector. SCORE-LM thus combines a compact radar specialist with a language interface for communicating fault-specific inspection guidance.
Figures & tables
Fig. 1: Rad-R capture setup, arranged from the authors’ dataset assets [ 1 ] . From left: front and side vehicle views, followed by the two sensor-plate views. The hardware includes the cascade radar, camera, inertial and environmental sensors, positioning, storage, and controlled-vibration hardware.
Fig. 2: Rad-R acquisition dashboard: yaw2, radar frame 2122, snapshot at 217 s [ 1 ] . Radar views, camera frames, and auxiliary-sensor traces illustrate the recorded context.
Fig. 3: Frozen SCORE signal pipeline, with the selective-state-space block and source self-supervision. Each peak contributes 512 temporal and 54 descriptor coordinates. Temporal and descriptor branches are fused before mean pooling over peaks. Mean-plus-last pooling over radar loops is distinct from frame or peak averaging. The three legacy descriptors are uncalibrated heuristics. The diagram is schematic; reconstruction is training-only. The language connection in Fig. 4 is added after the frozen diagnostic readout.
Fig. 4: Visual overview of score-to-language alignment. (a) Frozen SCORE and its directional readout produce four raw similarities; projected tokens and a question condition generation. (b) Three-layer nonlinear projector, with widths shown for Qwen. (c) Low-rank updates inside frozen decoder linears. (d) Source-only answer-token training updates the projector and LoRA, not SCORE or base-language weights. Tensor shapes and traces are schematic; no textual fault label is supplied as input.
Fig. 5: Ten-frame confusion matrices using a common row-percentage scale and five seeds. Each seed contributes 40 vibration, 40 misalignment, 38 blockage, and 40 degradation decisions. There is no true Healthy row; SCORE’s empty Healthy prediction column follows its four-fault directional readout.
Fig. 6: Nine-method low-label transfer with common source-window selections. Error bars are population SD across five resample means. SCORE additionally uses unlabeled source pretraining. One label is one ten-frame window. The 20-window budget uses seven folds because one fold has only 18 source blockage windows; other budgets use eight.
Method
1 frame
10 frames
50 frames
100 frames
10-frame BA
10-frame F1
RD-CNN ERM
69.68
71.94
75.62
80.00
72.12
71.36
CORAL (adapted)
68.94
70.18
72.50
73.75
70.29
68.85
MixStyle
61.99
63.29
65.00
73.75
63.34
63.45
DARM (adapted)
61.31
64.62
68.75
73.75
64.68
63.22
DDDG (adapted)
67.51
69.40
70.83
71.25
69.39
69.27
BDC (adapted)
59.00
63.69
67.71
67.50
63.72
64.48
TABLE I: Primary capture-mean recall by frame budget and pooled ten-frame balanced accuracy (BA) and macro-F1, in percent. All five seeds retained. Adapted methods use radar-specific implementations.
Fig. 7: Before and after SCORE encoding for 178 matched ten-frame windows: 160 source windows and 18 held-out Blockage1 windows. Healthy samples are source-only. The inputs have 566 coordinates and the embeddings 128. Both t-SNE fits use cosine distance, perplexity 30, 1,000 iterations, and identical random initialization; labels only set plot colors. Each panel has its own two-dimensional coordinates.
Input
1-frame recall
10-frame recall
10-frame F1
Full
84.56 (1.09)
88.39 (0.70)
88.20 (0.67)
Physical only
84.62 (1.60)
87.47 (1.47)
87.29 (1.37)
Temporal only
27.45 (1.17)
44.58 (4.85)
39.61 (5.32)
No legacy geometry
74.34 (3.24)
86.11 (2.25)
85.86 (2.26)
TABLE II: Component ablations: mean recall and macro-F1 (%). Full and Physical only use the five-seed directional protocol (population SD). The last two rows retain the earlier three-seed input-ablation protocol (sample SD).
Runtime
Graph
GPU median / p99 (ms)
DSP median (ms)
DSP + GPU median / p99 (ms)
Windows
DDDG
1.04 / 1.93
132.83
138.13 / 156.87
WSL
RadrNet-DS-CI
8.51 / 14.68
89.18
147.82 / 197.58
WSL
SCORE directional graph
9.97 / 15.03
79.69
92.86 / 122.19
TABLE III: Processing measurements from separate runs: batch one, one CPU thread, 20 warm-ups and 200 synchronized repeats. Composed timing is measured directly.
Fig. 8: Radar model size versus ten-frame fault accuracy, measured as mean capture recall over eight held-out recordings. Error bars: population SD across five seeds. Size counts active scalar inference weights and fitted readout coefficients; scaler statistics, running buffers, training-only heads, and all language modules are excluded. SCORE includes its 640 center/direction values; RadrNet excludes its unused severity head. Adapted methods and readout/supervision differences follow Sec. VI. The highlighted gain is relative to RadrNet-DS-CI, the highest-recall competitor.
Score transform
Projector
True fault
Decoder agreement
Softmax
Linear
20/25
22/25
Sigmoid
Linear
20/25
22/25
Raw
Linear
21/25
20/25
Sigmoid
Full MLP
6/25
5/25
Raw
Full MLP
23/25
24/25
TABLE IV: Historical projector development study, fixed 25 short diagnosis questions. The full MLP and linear projectors differ in capacity and initialization.
Measure
With rank-8 LoRA
LoRA removed
Correct fault, all questions
62/75 (82.7%)
45/75 (60.0%)
Decoder agreement, all
67/75 (89.3%)
47/75 (62.7%)
Correct fault, canonical windows
20/24
17/24
Decoder agreement, canonical
22/24
18/24
Output cap reached
0/75
20/75
Unsupported observations †
0/75
≥31/75
TABLE V: Current maintenance evaluation. Same scores, questions, jointly trained projector and decoding; only language LoRA is removed in the control. Counts cover 75 questions on 24 radar windows.
Fig. 9: Matched maintenance dialogue, Q005 (Blockage1, W11). Both models receive the same radar scores, question and jointly trained projector; the control removes language LoRA. The adapted answer is complete; the control is an exact excerpt. The signal icon is schematic. Highlighted boxes are post-hoc review annotations, not model outputs. Both fault labels are correct in this example, but the control asserts echo absence that the four-score input does not establish. Aggregate unsupported-observation counts exclude fault-label errors: the adapted model still makes 13/75 incorrect fault statements.
Backbone
Correct fault / 75
Agreement / 75
Unsupported / 75
Capped / 75
Time / answer
Qwen3.5-4B
62
67
0 †
0
4.75
Phi-4-mini-instruct
59
66
0 †
0
2.98
Ministral-3-3B-Instruct-2512
54
57
0 †
0
2.67
Granite-4.2-3B
65
70
0 †
0
3.88
† Answer-level assistant review of unsupported observations beyond the fault label.
TABLE VI: Matched maintenance QA with rank-8 language adapters and trained projectors. Greedy decoding, 384-token cap; 75 questions on 24 windows. Time is mean generation latency in seconds per answer on an RTX A4000, excluding model loading.