Organizations: School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore · School of Electronic Information, Central South University, Changsha 410083, China · College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences. In contrast to conventional pixel-domain visual understanding, IBFU conducts semantic analysis without fully decoding images into the pixel domain. Since pixel-level visual content is not explicitly reconstructed during inference, this paradigm reduces visual exposure within the processing pipeline and suits privacy-friendly Artificial Intelligence of Things (AIoT) applications. In this paper, we propose Bitstream Fine-grained Generator (BFG), a novel foundation model tailored for IBFU. BFG consists of two main components: a Bitstream Semantic Encoder (BSeE) and a Fine-grained Semantic Generator (FSeG). BSeE directly models semantic representations from encoded image bitstreams without explicit pixel reconstruction, while FSeG transforms the extracted bitstream semantics into detailed natural-language descriptions through autoregressive generation. To train BFG and comprehensively evaluate IBFU in practical AIoT scenarios, where image bitstreams may suffer corruption during transmission and storage, we construct a large-scale Corrupted-bitstream Fine-grained Understanding dataset (CFU-D), containing both intact bitstreams and corrupted variants across multiple corruption types and severity levels. Experiments show that BFG maintains stable fine-grained caption generation under bitstream corruption. For example, the performance only has slight change from 0.6339 to 0.6077 in terms of average CIDEr score on Stanford Dogs Caption dataset, while vision-language models, such as Qwen-VL-Chat, BLIP-2, GLM, Gemini, and GPT suffer severe performance decrease. This paper provides a practical paradigm for privacy-friendly fine-grained understanding in AIoT.
Figures & tables
Fig. 1 : Bitstream-domain fine-grained understanding framework for privacy-friendly AIoT.
Fig. 2 : Overall framework of BFG. BFG takes byte sequences as input and extracts bitstream features through ByteFormer, which consists of byte embedding, windowed Transformer encoding, and hierarchical sequence compression. The extracted features are then fed into a Transformer decoder with self-attention and cross-attention to autoregressively generate fine-grained semantic descriptions.
Fig. 3 : Illustration of corruptions on bitstream.
Fig. 4 : Confusion matrices of six bitstream-domain models on the 10 most frequently predicted Stanford Dogs breeds.
Model
Breed Acc
Precision
Recall
F1-score
Balanced Acc
BFG-Direct
58.51%
57.96%
58.06%
57.35%
58.06%
DCT-CNN
7.16%
5.49%
6.72%
4.35%
6.72%
BLOCK-CNN
19.02%
17.71%
18.51%
17.59%
18.51%
bGPT
1.38%
0.01%
0.83%
0.02%
0.83%
MBLM
4.08%
3.21%
4.06%
2.93%
4.06%
MegaByte
1.39%
0.01%
0.83%
0.02%
0.83%
TABLE I : Comparison of classification performance of bitstream-domain models.
Model
Strategy
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Breed Acc
Dec. Rate
BFG(ours)
No corrupt
0.6339
0.0996
0.3651
0.3408
0.3058
57.63%
-
Replace
0.6240
0.0966
0.3623
0.3375
0.3012
51.24%
-
Drop
0.6248
0.0970
0.3621
0.3376
0.3015
52.50%
-
Repeat
0.5479
0.0748
0.3310
0.3073
0.2645
20.59%
-
Average
0.6077
0.0920
0.3551
0.3308
0.2933
45.49%
-
Qwen-VL-Chat
No corrupt
0.4389
0.0806
0.2562
0.0000
0.2118
38.90%
100.00%
TABLE II : Performance comparison on Stanford Dogs Caption under different corruption strategies.
Model
Strategy
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Dec. Rate
BFG(ours)
No corrupt
0.0521
0.1063
0.3005
0.3055
0.8914
-
Replace
0.0516
0.1051
0.2993
0.3021
0.8915
-
Drop
0.0636
0.1064
0.3023
0.3046
0.8926
-
Repeat
0.0528
0.1062
0.2977
0.3053
0.8918
-
Average
0.0550
0.1060
0.2999
0.3044
0.8918
-
Qwen-VL-Chat
No corrupt
0.0043
0.0196
0.1812
0.2011
0.8634
100.00%
TABLE III : Performance comparison on Flickr30k under different corruption strategies.
Fig. 5 : CIDEr score comparison and retention rate heatmap on the Stanford Dogs Caption dataset under different corruption strategies: (a) CIDEr score comparison across models and corruption strategies; (b) CIDEr retention rate heatmap showing the percentage of baseline performance maintained.
Fig. 6 : CIDEr score comparison and retention rate heatmap on the Flickr30k dataset under different corruption strategies: (a) CIDEr score comparison across models and corruption strategies; (b) CIDEr retention rate heatmap showing the percentage of baseline performance maintained.
Fig. 7 : Robustness comparison of BFG and vision-language models on Flickr30k. Each axis represents one corruption condition. Top row is BERTScore-F1 and bottom row is METEOR. BFG covers the largest area across all conditions.
Fig. 8 : Visual comparison of model-generated descriptions under 2% byte replacement.
With byte masking
Without byte masking
Metric
Clean
Replace
Drop
Repeat
All corrupt
Clean
Replace
Drop
Repeat
All corrupt
CIDEr
0.6339
0.6240
0.6248
0.5479
0.5989
0.6473
0.6336
0.6347
0.5637
0.6107
BLEU-4
0.0996
0.0966
0.0970
0.0748
0.0895
0.1086
0.1028
0.1046
0.0720
0.0931
ROUGE-L
0.3651
0.3623
0.3621
0.3310
0.3518
0.3755
0.3680
0.3702
0.3262
0.3548
METEOR
0.3408
0.3375
0.3376
0.3073
0.3275
0.3506
0.3427
0.3453
0.3020
0.3300
Breed Acc.
57.63%
51.24%
52.50%
20.59%
41.44%
62.94%
55.02%
57.44%
20.18%
44.21%
TABLE IV : Ablation on byte-level corruption augmentation for Stanford Dogs Caption. Bold values indicate the smaller relative drop for each metric and corruption condition.
With byte masking
Without byte masking
Metric
Clean
Replace
Drop
Repeat
All corrupt
Clean
Replace
Drop
Repeat
All corrupt
CIDEr
0.0521
0.0537
0.0560
0.0548
0.0549
0.0498
0.0544
0.0543
0.0465
0.0517
BLEU-4
0.1063
0.1060
0.1056
0.1070
0.1062
0.1046
0.1030
0.1042
0.0924
0.0999
ROUGE-L
0.3005
0.3004
0.3021
0.2975
0.3000
0.2956
0.2950
0.2963
0.2860
0.2925
METEOR
0.3055
0.3034
0.3031
0.3078
0.3048
0.3019
0.3003
0.3007
0.2983
0.2998
BERTScore-F1
0.8914
0.8915
0.8926
0.8918
0.8920
0.8900
0.8897
0.8899
0.8868
0.8888
TABLE V : Ablation on byte-level corruption augmentation for Flickr30k. Bold values indicate the smaller relative drop for each metric and corruption condition.
Format
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Breed Acc.
JPEG
0.6339
0.0996
0.3651
0.3408
0.3058
57.63%
WebP
0.4077
0.0488
0.2913
0.2748
0.1944
0.44%
PNG
0.3381
0.0526
0.2934
0.2787
0.1964
0.44%
TABLE VI : Zero-shot evaluation on additional image formats.
Fig. 9 : Visual comparison of model-generated descriptions on Flickr30k different bitstream corruption.
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China · School of Information Science and Engineering, Linyi University, Linyi 276000, China
Hefei University of Technology · Key Laboratory of Knowledge Engineering with Big Data · Key Laboratory of Knowledge Engineering with Big Data (Hefei University of Technology), Ministry of Education; School of Computer Science and Information Engineering, Hefei University of Technology; and Intelligent Interconnected Systems Laboratory of Anhui Province (Hefei University of Technology) +3