Instruction-Conditioned Electromagnetic Spectrum Understanding via Budget-Adaptive Signal Tokenization
Organizations: School of Artificial Intelligence, Xidian University, Xi’an, China
Abstract
Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection. Multimodal large language models offer a unified interface, but extending vision-language models (VLMs) to raw I/Q signals requires tokenization that balances fidelity against a strict budget. For signals, dense encoding causes token costs to grow with observation length, whereas fixed-resolution compression may discard short-duration or localized signal evidence. Thus, we propose \textbf{BATok}, a budget-adaptive signal tokenizer that adjusts token capacity to the input length while allocating that capacity according to the signal content. BATok constructs candidate representations from signal-derived features using lightweight multi-resolution branches, then combines a local energy prior with learnable queries to resample these representations into compact signal tokens. The number of tokens adapts to the input length while remaining strictly bounded. The resulting tokens are projected into the language embedding space of VLMs. We further introduce \textbf{EMSpec-Instruct}, a multimodal instruction dataset aligning raw I/Q signals, waterfall images, and language supervision for modulation recognition, structured detection, and language-conditioned signal grounding. Experiments show that BATok learns effective signal representations and achieves competitive performance across all tasks.
Figures & tables
| Modulation Recognition | Structured Detection | Signal Grounding | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Backbone | BATok | OA (%) | Macro-F1 (%) | Parse (%) | [email protected] (%) | mIoU (%) | Parse (%) | [email protected] (%) | mIoU (%) | Parse (%) | Empty Acc. (%) |
| Qwen3-VL-2B | 33.90 | 35.38 | 100.00 | 62.25 | 94.66 | 98.08 | 82.52 | 94.75 | 99.99 | 93.19 | |
| Qwen3-VL-2B | ✓ | 40.91 | 41.26 | 86.02 | 64.27 | 95.20 | 98.29 | 83.00 | 95.24 | 99.96 | 92.43 |
| LLaVA-1.5-7B | 25.58 | 26.26 | 100.00 | 45.99 | 88.72 | 99.43 | 66.94 | 88.81 | 99.95 | 91.92 | |
| LLaVA-1.5-7B | ✓ | 44.14 | 43.20 | 94.95 | 47.80 | 89.40 | 97.17 | 68.27 | 89.39 | 99.87 | 91.23 |
| Method | Acc. (%) | Macro-F1 (%) |
|---|---|---|
| MAMC ( Zhang et al. 2024 ) | 43.20 | 42.31 |
| SiT ( Zhai et al. 2025 ) | 53.56 | 53.36 |
| DNCNet ( Du et al. 2022 ) | 58.18 | 56.89 |
| BATok | 58.64 | 58.54 |
| Group | Variant | Acc. (%) | Macro-F1 (%) |
|---|---|---|---|
| Token budget | Fixed | 57.68 | 57.53 |
| Fixed | 57.97 | 57.60 | |
| Linear-high | 58.70 | 58.48 | |
| Piecewise | 58.64 | 58.54 | |
| Input features | I/Q only | 56.37 | 56.30 |
| I/Q + + + | 58.64 | 58.54 |
| Branch | Kernel (samples) | Stride (samples) | Approximate length (samples) |
|---|---|---|---|
| Small | 7 | 4 | |
| Mid | 31 | 16 | |
| Large | 127 | 64 |
| Valid input length (samples) | Total budget (tokens) |
|---|---|
| 32 | |
| 48 | |
| 64 | |
| 96 | |
| 128 | |
| 192 |
| Setting | Qwen3-VL-2B | LLaVA-1.5-7B |
|---|---|---|
| LLM adaptation | LoRA ( ) | LoRA ( ) |
| Vision tower | Trainable | Trainable |
| Native multimodal projector | Trainable | Trainable |
| BATok | Trainable | Trainable |
| Signal projector | Trainable | Trainable |
| Signal position handling | Extended M-RoPE | Learned 1D embedding |
| Source | Task stream | Category coverage | Train (observations) | Test (observations) |
|---|---|---|---|---|
| HisarMod-Sub | Recognition | 26 classes: PSK 6 , QAM 7 , FSK 4 , PAM 3 , and six analog AM/FM/PM variants | 44,000 | 22,000 |
| EMSpec-Sim-Sub | Recognition | 34 classes: generic digital/analog modulations, WiFi 6 , BLE LE1M/LE2M, LoRa 125/250/500 kHz, Zigbee, and SRRC-shaped QPSK/16QAM | 22,000 | 11,000 |
| SpaceNet | Det./ground. | 11 wireless classes: WiFi 6 , BLE LE1M/LE2M, LoRa 250 kHz, Zigbee, and SRRC-shaped QPSK | 7,500 | 2,500 |
| Multi-signal simulation | Det./ground. | 14 classes: the SpaceNet vocabulary augmented with AM, FM, and SRRC-shaped 16QAM | 20,000 | 7,500 |
| Total | 93,500 | 43,000 |
| Task | Example instruction | Target |
|---|---|---|
| Recognition | Classify the modulation using both modalities. | <c>["8ASK"]</c> |
| Detection | Detect every signal and return all boxes and classes. | <d><b>[[339,166,982,833]]</b> <c>["16QAM"]</c></d> |
| Grounding | Find regions that overlap 2419.589–2448.466 MHz. | <b>[[339,166,982,833],...]</b> |
| Negative grounding | Find all BLE LE2M regions in this scene. | <b>[]</b> |
| Task | Train (records) | Train (%) | Test (records) | Test (%) |
|---|---|---|---|---|
| Modulation recognition | 120,000 | 30.00 | 33,000 | 48.61 |
| Structured detection | 27,500 | 6.88 | 10,000 | 14.73 |
| Signal grounding | 252,500 | 63.12 | 24,883 | 36.66 |
| Total | 400,000 | 100.00 | 67,883 | 100.00 |
| Grounding subtype | Train (records) | Test (records) |
|---|---|---|
| Class | 27,500 | 1,631 |
| Class fallback (no overlap) | – | 313 |
| Time | 27,500 | 1,656 |
| Frequency | 27,500 | 1,640 |
| Class–time–frequency | 27,500 | 1,656 |
| Longest duration | 27,500 | 1,638 |
| Mod. Rec. | Detection | Grounding | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Backbone | Test Input | OA (%) | Macro-F1 (%) | Parse (%) | [email protected] (%) | mIoU (%) | Parse (%) | [email protected] (%) | mIoU (%) | Parse (%) | Empty Acc. (%) |
| Qwen3-VL-2B + BATok | Image + Text | 16.52 | 11.93 | 100.00 | 64.12 | 95.20 | 97.88 | 81.78 | 95.23 | 99.94 | 92.04 |
| Qwen3-VL-2B + BATok | Signal + Text | 40.36 | 40.33 | 88.96 | 3.64 | 72.48 | 99.44 | 13.36 | 74.37 | 99.90 | 89.39 |
| Qwen3-VL-2B + BATok | Image + Signal + Text | 40.91 | 41.26 | 86.02 | 64.27 | 95.20 | 98.29 | 83.00 | 95.24 | 99.96 | 92.43 |
| LLaVA-1.5-7B + BATok | Image + Text | 3.08 | 2.54 | 99.35 | 47.69 | 89.28 | 98.19 | 67.75 | 89.29 | 99.88 | 85.15 |
| LLaVA-1.5-7B + BATok | Signal + Text | 43.68 | 43.02 | 94.90 | 3.06 | 74.68 | 98.80 | 12.85 | 75.05 | 99.81 | 89.11 |
| Training tasks | Grounding | |||
| Mod. Rec. | Detection | Grounding | [email protected] (%) | mIoU (%) |
| ✓ | 81.97 | 95.15 | ||
| ✓ | ✓ | 81.99 | 95.24 | |
| ✓ | ✓ | 81.47 | 95.06 | |
| ✓ | ✓ | ✓ | 83.00 | 95.24 |
| Method | Input | Acc. (%) | Macro-F1 (%) | Params (M) | Latency (ms/sample) |
|---|---|---|---|---|---|
| ResNet-50 | Waterfall | 32.43 | 34.33 | 23.610 | 4.105 |
| ViT | Waterfall | 31.06 | 33.13 | 85.837 | 3.424 |
| MAMC ( Zhang et al. 2024 ) | I/Q | 29.62 | 27.64 | 0.009 | 1.466 |
| SiT ( Zhai et al. 2025 ) | I/Q | 41.56 | 41.43 | 0.654 | 0.967 |
| DNCNet ( Du et al. 2022 ) | I/Q | 44.17 | 42.80 | 3.870 | 1.860 |
| BATok | I/Q | 49.36 | 47.19 | 1.047 | 5.443 |
| Method | mIoU (ratio) | [email protected] (ratio) | [email protected] (ratio) | [email protected] (ratio) |
|---|---|---|---|---|
| YOLO26n | 0.9291 | 0.6466 | 0.6288 | 0.6654 |
| YOLO26l | 0.9475 | 0.6865 | 0.6514 | 0.7255 |
| YOLO26x | 0.9517 | 0.6925 | 0.6577 | 0.7311 |
| DETR | 0.8597 | 0.4127 | 0.2918 | 0.7044 |
| Backbone | Detection measure | Image + Text (%) | Image + Signal + Text (%) |
|---|---|---|---|
| Qwen3-VL-2B | Class-aware [email protected] | 62.25 | 64.27 |
| Class-agnostic [email protected] | 90.74 | 90.68 | |
| Conditional class accuracy | 68.57 | 70.72 | |
| Class-agnostic matched-box mIoU | 95.02 | 95.48 | |
| Class-agnostic [email protected] | 82.87 | 82.66 | |
| Class-agnostic [email protected]:0.95 | 71.23 | 72.19 |
| Qwen3-VL-2B + BATok | LLaVA-1.5-7B + BATok | |||||||
|---|---|---|---|---|---|---|---|---|
| Query Type | [email protected] (%) | mIoU (%) | Empty Acc. (%) | Parse (%) | [email protected] (%) | mIoU (%) | Empty Acc. (%) | Parse (%) |
| Modulation-conditioned | 83.51 | 94.85 | – | 99.94 | 66.59 | 87.95 | – | 100.00 |
| Time-conditioned | 90.28 | 95.22 | – | 99.88 | 74.07 | 89.24 | – | 99.58 |
| Frequency-conditioned | 87.02 | 94.82 | – | 100.00 | 74.16 | 88.91 | – | 99.88 |
| Compositional | 89.46 | 94.50 | – | 99.94 | 79.49 | 88.80 | – | 99.88 |
| Overlap / relational | 78.00 | 95.58 | – | 99.92 | 66.13 | 90.01 | – | 99.87 |
| HisarMod2019 | EMSpec-Sim | |||||
|---|---|---|---|---|---|---|
| Method | Acc. (%) | Macro-F1 (%) | Acc. (%) | Macro-F1 (%) | Params (M) | Latency (ms/sample) |
| MAMC ( Zhang et al. 2024 ) | 55.22 | 53.23 | 31.18 | 31.39 | 0.008 | 1.470 |
| SiT ( Zhai et al. 2025 ) | 71.34 | 71.30 | 35.78 | 35.41 | 0.649 | 0.978 |
| DNCNet ( Du et al. 2022 ) | 75.25 | 75.19 | 41.11 | 38.58 | 3.856 | 1.868 |
| BATok | 77.53 | 77.53 | 39.75 | 39.54 | 1.044 | 5.509 |
| Strategy | Avg. (tokens) | Max (tokens) | Avg. Acc. (%) | Avg. F1 (%) | Latency (ms/sample) |
|---|---|---|---|---|---|
| Fixed-16 | 16.00 | 16 | 57.91 | 57.62 | 5.358 |
| Fixed-32 | 32.00 | 32 | 57.68 | 57.53 | 5.294 |
| Fixed-48 | 48.00 | 48 | 57.97 | 57.60 | 5.300 |
| Fixed-64 | 64.00 | 64 | 58.04 | 57.70 | 5.310 |
| Fixed-128 | 128.00 | 128 | 58.08 | 57.90 | 5.469 |
| Linear-low | 32.50 | 48 | 58.60 | 58.34 | 5.469 |
| Strategy | Avg. (tokens) | (samples) Acc./F1 (%) | (samples) Acc./F1 (%) | (samples) Acc./F1 (%) | (samples) Acc./F1 (%) | (samples) Acc./F1 (%) | (samples) Acc./F1 (%) |
|---|---|---|---|---|---|---|---|
| Fixed-32 | 32.00 | 31.99 /30.56 | 34.69/33.34 | 38.72 /36.74 | 41.28/39.43 | 40.39 /38.67 | 40.76/39.24 |
| Fixed-39 | 39.00 | 31.73/ 30.77 | 36.19/ 35.39 | 38.10/ 37.21 | 41.40/39.97 | 38.93/37.96 | 41.18/ 40.70 |
| Fixed-48 | 48.00 | 29.23/28.76 | 34.27/34.04 | 35.05/34.77 | 40.62/ 40.35 | 39.54/ 39.25 | 39.86/39.23 |
| Piecewise | 39.23 | 31.84/29.70 | 36.73 /34.54 | 37.35/35.44 | 42.11 /38.93 | 40.39 /37.97 | 41.42 /38.99 |
| Group | Variant | HisarMod2019 Acc./F1 (%) | EMSpec-Sim Acc./F1 (%) | Avg. Acc. (%) | Avg. F1 (%) |
| Input features | I/Q only | 74.07/73.97 | 38.66/38.62 | 56.37 | 56.30 |
| I/Q+ + + | 77.53/77.53 | 39.75/39.54 | 58.64 | 58.54 | |
| Temporal scales | Small | 68.89/68.88 | 40.51/40.18 | 54.70 | 54.53 |
| Small+mid | 72.64/72.53 | 39.48/39.63 | 56.06 | 56.08 | |
| Small+mid+large | 77.53/77.53 | 39.75/39.54 | 58.64 | 58.54 | |
| Resampling bias | Energy prior only | 76.99/77.07 | 36.95/37.36 | 56.97 | 57.22 |
| GT | LLaVA | LLaVA+BATok | Qwen3-VL | Qwen3-VL+BATok |
|---|---|---|---|---|
| GT | LLaVA | LLaVA+BATok | Qwen3-VL | Qwen3-VL+BATok |
|---|---|---|---|---|
| GT | LLaVA | LLaVA+BATok | Qwen3-VL | Qwen3-VL+BATok |
| (a) Class-conditioned query | ||||
| (b) Frequency-conditioned query | ||||
| GT | LLaVA | LLaVA+BATok | Qwen3-VL | Qwen3-VL+BATok |
| (a) Time-conditioned query | ||||
| (b) Overlap-conditioned query | ||||