Electromagnetic spectrum monitoring increasingly requires flexible analysis beyond task-specific recognition and detection. Multimodal large language models offer a unified interface, but extending vision-language models (VLMs) to raw I/Q signals requires tokenization that balances fidelity against a strict budget. For signals, dense encoding causes token costs to grow with observation length, whereas fixed-resolution compression may discard short-duration or localized signal evidence. Thus, we propose \textbf{BATok}, a budget-adaptive signal tokenizer that adjusts token capacity to the input length while allocating that capacity according to the signal content. BATok constructs candidate representations from signal-derived features using lightweight multi-resolution branches, then combines a local energy prior with learnable queries to resample these representations into compact signal tokens. The number of tokens adapts to the input length while remaining strictly bounded. The resulting tokens are projected into the language embedding space of VLMs. We further introduce \textbf{EMSpec-Instruct}, a multimodal instruction dataset aligning raw I/Q signals, waterfall images, and language supervision for modulation recognition, structured detection, and language-conditioned signal grounding. Experiments show that BATok learns effective signal representations and achieves competitive performance across all tasks.
Figures & tables
Figure 1: Motivation for unified instruction-conditioned spectrum understanding. Task-specific models require manual redesign, while detector–LLM cascades inherit upstream errors. Our framework couples signal tokenization with a VLM through a unified interface.
Figure 2: Overview of BATok and its integration with a VLM. Variable-length I/Q signals are augmented with signal-derived features, encoded at multiple temporal resolutions, assigned a bounded length-adaptive token budget, and resampled using learned relevance and an energy prior. The resulting multi-resolution signal tokens are aggregated, projected into the language embedding space, and jointly processed with image and text tokens for unified spectrum understanding.
Figure 3: Construction and statistics of EMSpec-Instruct. Structured observation- and instance-level metadata are converted into metadata-grounded instructions. The bottom panels summarize the task composition of the training and test sets and the distribution of signal lengths.
Table 1: Unified spectrum-understanding results on EMSpec-Instruct. The original VLMs receive image–text inputs, whereas their BATok-enhanced counterparts receive aligned image–signal–text inputs. Parse is the proportion of outputs satisfying the required task-specific format, and Empty Acc. is the accuracy of returning an empty set for unmatched grounding queries. Within each backbone, the better value between the original VLM and its BATok-enhanced counterpart is bolded.
Figure 4: Visualization of detection and grounding results. Detection question: “Enumerate every transmission using aligned box and class arrays.” Ground truth: [141,424,846,924] (WiFi 20 MHz 64QAM) and [145,443,842,943] (WiFi 20 MHz QPSK). Grounding question: “Return boxes of the longest-duration transmissions.” Ground truth: [2,956,887,969] .
Method
Acc. (%) ↑
Macro-F1 (%) ↑
MAMC ( Zhang et al. 2024 )
43.20
42.31
SiT ( Zhai et al. 2025 )
53.56
53.36
DNCNet ( Du et al. 2022 )
58.18
56.89
BATok
58.64
58.54
Table 2: Standalone raw-I/Q modulation-recognition results. Each average is the unweighted mean across the two primary benchmarks, HisarMod2019 and EMSpec-Sim . Best and second-best averages are bolded and underlined, respectively. More results are provided in Supplementary Section S4.1.
Group
Variant
Acc. (%) ↑
Macro-F1 (%) ↑
Token budget
Fixed K=32
57.68
57.53
Fixed K=48
57.97
57.60
Linear-high
58.70
58.48
Piecewise K(T)†
58.64
58.54
Input features
I/Q only
56.37
56.30
I/Q + A + P + Δϕ†
58.64
58.54
Table 3: Ablation study of BATok in standalone modulation recognition. Accuracy and Macro-F1 are percentages averaged over HisarMod2019 and EMSpec-Sim ; † denotes the full BATok configuration within each ablation group.
Branch
Kernel (samples)
Stride (samples)
Approximate length (samples)
Small
7
4
⌈T/4⌉
Mid
31
16
⌈T/16⌉
Large
127
64
⌈T/64⌉
Table S1: Temporal-branch configuration used by BATok. The output length is approximate because convolution padding and the valid-length mask are handled separately.
Valid input length T (samples)
Total budget K(T) (tokens)
T≤212
32
212<T≤215
48
215<T≤216
64
216<T≤218
96
218<T≤220
128
220<T≤222
192
Table S2: BATok token schedule.
Setting
Qwen3-VL-2B
LLaVA-1.5-7B
LLM adaptation
LoRA ( r=32 )
LoRA ( r=8 )
Vision tower
Trainable
Trainable
Native multimodal projector
Trainable
Trainable
BATok
Trainable
Trainable
Signal projector
Trainable
Trainable
Signal position handling
Extended M-RoPE
Learned 1D embedding
Table S3: Backbone-specific VLM integration settings. The language models are adapted with LoRA, while BATok and its projector are trained and saved as complete modules.
Source
Task stream
Category coverage
Train (observations)
Test (observations)
HisarMod-Sub
Recognition
26 classes: PSK 6 , QAM 7 , FSK 4 , PAM 3 , and six analog AM/FM/PM variants
44,000
22,000
EMSpec-Sim-Sub
Recognition
34 classes: generic digital/analog modulations, WiFi 6 , BLE LE1M/LE2M, LoRa 125/250/500 kHz, Zigbee, and SRRC-shaped QPSK/16QAM
22,000
11,000
SpaceNet
Det./ground.
11 wireless classes: WiFi 6 , BLE LE1M/LE2M, LoRa 250 kHz, Zigbee, and SRRC-shaped QPSK
7,500
2,500
Multi-signal simulation
Det./ground.
14 classes: the SpaceNet vocabulary augmented with AM, FM, and SRRC-shaped 16QAM
20,000
7,500
Total
93,500
43,000
Table S4: Source observations and category coverage used to construct EMSpec-Instruct. Counts refer to electromagnetic observations before instruction expansion. WiFi 6 denotes the six combinations of 20/40-MHz bandwidth and QPSK/16QAM/64QAM modulation.
Task
Example instruction
Target
Recognition
Classify the modulation using both modalities.
<c>["8ASK"]</c>
Detection
Detect every signal and return all boxes and classes.
Table S5: Representative task realizations and target formats. Scene context is abbreviated in the table but is present in the actual time/frequency-conditioned prompts.
Task
Train (records)
Train (%)
Test (records)
Test (%)
Modulation recognition
120,000
30.00
33,000
48.61
Structured detection
27,500
6.88
10,000
14.73
Signal grounding
252,500
63.12
24,883
36.66
Total
400,000
100.00
67,883
100.00
Table S6: Final task composition. The validation set is a balanced 3,000-example subset sampled from the held-out task files and is not included in the test total.
Grounding subtype
Train (records)
Test (records)
Class
27,500
1,631
Class fallback (no overlap)
–
313
Time
27,500
1,656
Frequency
27,500
1,640
Class–time–frequency
27,500
1,656
Longest duration
27,500
1,638
Table S7: Signal-grounding subtype counts. A dash denotes a subtype that is not selected for that partition.
Table S8: BATok-enhanced VLM performance by input configuration. Values are percentages, and all input configurations for a backbone use the same trained checkpoint. Parse denotes successful task-format parsing, and Empty Acc. is reported for unmatched grounding queries.
Table S9: Grounding performance under different instruction-tuning task compositions. All variants use Qwen3-VL-2B+BATok and are evaluated on the same grounding test set.
Method
Input
Acc. (%) ↑
Macro-F1 (%) ↑
Params (M) ↓
Latency (ms/sample) ↓
ResNet-50
Waterfall
32.43
34.33
23.610
4.105
ViT
Waterfall
31.06
33.13
85.837
3.424
MAMC ( Zhang et al. 2024 )
I/Q
29.62
27.64
0.009
1.466
SiT ( Zhai et al. 2025 )
I/Q
41.56
41.43
0.654
0.967
DNCNet ( Du et al. 2022 )
I/Q
44.17
42.80
3.870
1.860
BATok
I/Q
49.36
47.19
1.047
5.443
Table S10: Waterfall-image and raw-I/Q modulation recognition on the aligned EMSpec-Instruct subset, defined as the union of HisarMod-Sub and EMSpec-Sim-Sub . Latency is measured in milliseconds per sample with batch size 1 on one NVIDIA GeForce RTX 4090.
Table S11: Task-specific detection results on the EMSpec-Instruct detection test set. Best and second-best values are bolded and underlined, respectively. These systems use dedicated detection heads.
Table S12: Decomposition of structured-detection gains into class semantics and box geometry. Class-agnostic metrics ignore the predicted class during IoU matching; conditional class accuracy is computed only over boxes matched at IoU ≥0.5 .
Table S13: Language-conditioned grounding results by query type for the image–signal–text test configuration. Values are percentages; a dash marks an inapplicable metric, and Parse denotes successful required-format parsing. Overlap/relational pools duration, bandwidth, brightness, and overlap queries, while unmatched pools negative-class and class-fallback-no-overlap queries.
HisarMod2019
EMSpec-Sim
Method
Acc. (%) ↑
Macro-F1 (%) ↑
Acc. (%) ↑
Macro-F1 (%) ↑
Params (M) ↓
Latency (ms/sample) ↓
MAMC ( Zhang et al. 2024 )
55.22
53.23
31.18
31.39
0.008
1.470
SiT ( Zhai et al. 2025 )
71.34
71.30
35.78
35.41
0.649
0.978
DNCNet ( Du et al. 2022 )
75.25
75.19
41.11
38.58
3.856
1.868
BATok
77.53
77.53
39.75
39.54
1.044
5.509
Table S14: Per-benchmark standalone raw-I/Q modulation-recognition results and classification efficiency. Best and second-best performance values within each benchmark are bolded and underlined, respectively. Latency is measured with batch size 1 on one NVIDIA GeForce RTX 4090.
Strategy
Avg. K (tokens)
Max K (tokens)
Avg. Acc. (%)
Avg. F1 (%)
Latency (ms/sample)
Fixed-16
16.00
16
57.91
57.62
5.358
Fixed-32
32.00
32
57.68
57.53
5.294
Fixed-48
48.00
48
57.97
57.60
5.300
Fixed-64
64.00
64
58.04
57.70
5.310
Fixed-128
128.00
128
58.08
57.90
5.469
Linear-low
32.50
48
58.60
58.34
5.469
Table S15: Unified performance–budget–latency ablation. Mean Accuracy and Macro-F1 are the unweighted means across HisarMod2019 and EMSpec-Sim . Latency is measured in milliseconds per sample on one NVIDIA GeForce RTX 4090. † denotes the full BATok configuration. Best and second-best mean results are bolded and underlined, respectively.
Strategy
Avg. K (tokens)
T=1,024 (samples) Acc./F1 (%)
T=2,048 (samples) Acc./F1 (%)
T=4,096 (samples) Acc./F1 (%)
T=8,192 (samples) Acc./F1 (%)
T=16,384 (samples) Acc./F1 (%)
T=32,768 (samples) Acc./F1 (%)
Fixed-32
32.00
31.99 /30.56
34.69/33.34
38.72 /36.74
41.28/39.43
40.39 /38.67
40.76/39.24
Fixed-39
39.00
31.73/ 30.77
36.19/ 35.39
38.10/ 37.21
41.40/39.97
38.93/37.96
41.18/ 40.70
Fixed-48
48.00
29.23/28.76
34.27/34.04
35.05/34.77
40.62/ 40.35
39.54/ 39.25
39.86/39.23
Piecewise K(T)†
39.23
31.84/29.70
36.73 /34.54
37.35/35.44
42.11 /38.93
40.39 /37.97
41.42 /38.99
Table S16: Controlled within-dataset token-budget comparison. EMSpec-Sim-VarLen is a variable-length subset of EMSpec-Sim . Each length column reports Accuracy/Macro-F1 in percent; Avg. K is weighted by the number of test samples, and bold marks the best result for each metric at each length.
Group
Variant
HisarMod2019 Acc./F1 (%)
EMSpec-Sim Acc./F1 (%)
Avg. Acc. (%)
Avg. F1 (%)
Input features
I/Q only
74.07/73.97
38.66/38.62
56.37
56.30
I/Q+ A + P + Δϕ†
77.53/77.53
39.75/39.54
58.64
58.54
Temporal scales
Small
68.89/68.88
40.51/40.18
54.70
54.53
Small+mid
72.64/72.53
39.48/39.63
56.06
56.08
Small+mid+large †
77.53/77.53
39.75/39.54
58.64
58.54
Resampling bias
Energy prior only
76.99/77.07
36.95/37.36
56.97
57.22
Table S17: Per-dataset component ablations for standalone BATok classification. Dataset columns report Accuracy/Macro-F1 in percent; Avg. columns are arithmetic means across HisarMod2019 and EMSpec-Sim . † denotes the full BATok configuration, and bold marks the best cross-dataset mean within each ablation group.
GT
LLaVA
LLaVA+BATok
Qwen3-VL
Qwen3-VL+BATok
Figure S1: Detection of a broadband transmission and a narrow-band target. Question: “Detect all occupied signal regions and pair each bounding box with the matching class label.” Ground truth: [285,500,904,1000] (WiFi 20 MHz 64QAM) and [346,843,958,849] (LoRa 250 kHz). Coordinates are normalized to [0,1000] ; the GT panel uses red boxes.
GT
LLaVA
LLaVA+BATok
Qwen3-VL
Qwen3-VL+BATok
Figure S2: Detection in a crowded heterogeneous scene. Question: “Perform exhaustive detection and respond only with the requested tags.” Ground truth: [289,25,888,691] (WiFi 20 MHz 64QAM), [304,929,897,962] (BLE LE1M), [308,392,863,459] (SRRC QPSK), and [310,261,828,928] (WiFi 20 MHz QPSK). Coordinates are normalized to [0,1000] ; the GT panel uses red boxes.
GT
LLaVA
LLaVA+BATok
Qwen3-VL
Qwen3-VL+BATok
(a) Class-conditioned query
(b) Frequency-conditioned query
Figure S3: Grounding under class and frequency constraints. (a) Question: “Which regions correspond to the LoRa 250 kHz class?” GT: [46,123,741,129] . (b) Question: “Find signal regions that overlap the frequency band 2455.015597–2457.837200 MHz.” GT: [77,0,1000,800] and [96,230,994,270] . Coordinates are normalized to [0,1000] ; GT panels use red boxes.
GT
LLaVA
LLaVA+BATok
Qwen3-VL
Qwen3-VL+BATok
(a) Time-conditioned query
(b) Overlap-conditioned query
Figure S4: Grounding under time and overlap constraints. (a) Question: “Find signal regions that overlap the time interval 1.092–10.124 ms.” GT: [46,891,483,916] and [96,886,463,911] . (b) Question: “Find the signal regions involved in time–frequency overlap.” GT: [152,0,886,400] and [788,0,1000,800] . The overlap case is deliberately included because all compared variants succeed on it. Coordinates are normalized to [0,1000] ; GT panels use red boxes.
This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids. It introduces a preliminary continuous-token model built around a one-level Haar DWT/IDWT frontend, a shared coefficient-token layout, optional structural metadata, lightweight modality value adapters, and a shared token-wise encoder-decoder trunk. On Speech Commands, EuroSAT RGB, and DAVIS 2017 data, a dense shared model reaches 39.92 dB audio, 29.37 dB image, and 23.93 dB video PSNR. A matched-rate sweep under continuous latent scalar budgets indicates that the visual gains are not explained solely by latent capacity, while also showing that additive metadata embeddings are not a universal source of improvement. Finally, fixed-rate energy selection provides a strong non-parametric baseline: energy_global improves average PSNR over uniform selection by 16.73 dB for audio, 16.90 dB for images, and 15.86 dB for video under compressed keep ratios. Masked sparse training reaches 34.45 dB video PSNR with 50% of dense tokens. The results support a unified wavelet token schema and sparse token interface, while stopping short of establishing a universal discrete vocabulary.
Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose M2Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the M2Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.
Chunpu Xu, Zhixuan Liang, Yuhao Zhang +6
The Hong Kong Polytechnic University, HongKong SAR, China · Shanghai AI Laboratory, Shanghai, China · The University of Hong Kong, HongKong SAR, China +2
Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory and computational resources. Post-training quantization (PTQ) has successfully compressed language models to as low as 1-bit precision, its effectiveness for multimodal LLMs (MLLMs) remains unexplored. In this paper, we present the first method for ultra-low-bit (<4-bit) quantization of MLLMs. Our analysis reveals that multimodal tokens and intermediate layer activations produced by them exhibit significantly higher entropy compared to text tokens, indicating greater functional complexity that makes MLLMs less tolerant to ultra-low bit quantization. However, this entropy varies significantly across layers, with some layers producing lower-entropy activation distributions that we empirically show can better tolerate ultra-low bit quantization. Existing PTQ methods optimize weight quantization within each layer but apply the same target precision uniformly, ignoring this variation in complexity across layers. Building on this insight, we propose LUQ: Layerwise Ultra-Low Bit Quantization, which characterizes each transformer layer's functional complexity via its output activation entropy and selectively applies ultra-low bit quantization to layers encoding simpler, more compressible functions. We also show that multimodal calibration (image and text tokens) boosts VQA performance in the ultra-low bit regime. Evaluated on LLaVA-1.5 and Qwen-2.5-VL across 9 VQA benchmarks, LUQ models use 40% and 31% less memory than their 4-bit counterparts while exhibiting less than 10% degradation on MME.
Shubhang Bhatnagar, Andy Xu, Kar-Han Tan +1
University of Illinois Urbana-Champaign · Work done as an intern at HP Inc. · University of California, Los Angeles +1