Organizations: School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China · School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore · School of Electronic Information, Central South University, Changsha 410083, China · College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences. In contrast to conventional pixel-domain visual understanding, IBFU conducts semantic analysis without fully decoding images into the pixel domain. Since pixel-level visual content is not explicitly reconstructed during inference, this paradigm reduces visual exposure within the processing pipeline and suits privacy-friendly Artificial Intelligence of Things (AIoT) applications. In this paper, we propose Bitstream Fine-grained Generator (BFG), a novel foundation model tailored for IBFU. BFG consists of two main components: a Bitstream Semantic Encoder (BSeE) and a Fine-grained Semantic Generator (FSeG). BSeE directly models semantic representations from encoded image bitstreams without explicit pixel reconstruction, while FSeG transforms the extracted bitstream semantics into detailed natural-language descriptions through autoregressive generation. To train BFG and comprehensively evaluate IBFU in practical AIoT scenarios, where image bitstreams may suffer corruption during transmission and storage, we construct a large-scale Corrupted-bitstream Fine-grained Understanding dataset (CFU-D), containing both intact bitstreams and corrupted variants across multiple corruption types and severity levels. Experiments show that BFG maintains stable fine-grained caption generation under bitstream corruption. For example, the performance only has slight change from 0.6339 to 0.6077 in terms of average CIDEr score on Stanford Dogs Caption dataset, while vision-language models, such as Qwen-VL-Chat, BLIP-2, GLM, Gemini, and GPT suffer severe performance decrease. This paper provides a practical paradigm for privacy-friendly fine-grained understanding in AIoT.
Figures & tables
Fig. 1 : Bitstream-domain fine-grained understanding framework for privacy-friendly AIoT.
Fig. 2 : Overall framework of BFG. BFG takes byte sequences as input and extracts bitstream features through ByteFormer, which consists of byte embedding, windowed Transformer encoding, and hierarchical sequence compression. The extracted features are then fed into a Transformer decoder with self-attention and cross-attention to autoregressively generate fine-grained semantic descriptions.
Fig. 3 : Illustration of corruptions on bitstream.
Fig. 4 : Confusion matrices of six bitstream-domain models on the 10 most frequently predicted Stanford Dogs breeds.
Model
Breed Acc
Precision
Recall
F1-score
Balanced Acc
BFG-Direct
58.51%
57.96%
58.06%
57.35%
58.06%
DCT-CNN
7.16%
5.49%
6.72%
4.35%
6.72%
BLOCK-CNN
19.02%
17.71%
18.51%
17.59%
18.51%
bGPT
1.38%
0.01%
0.83%
0.02%
0.83%
MBLM
4.08%
3.21%
4.06%
2.93%
4.06%
MegaByte
1.39%
0.01%
0.83%
0.02%
0.83%
TABLE I : Comparison of classification performance of bitstream-domain models.
Model
Strategy
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Breed Acc
Dec. Rate
BFG(ours)
No corrupt
0.6339
0.0996
0.3651
0.3408
0.3058
57.63%
-
Replace
0.6240
0.0966
0.3623
0.3375
0.3012
51.24%
-
Drop
0.6248
0.0970
0.3621
0.3376
0.3015
52.50%
-
Repeat
0.5479
0.0748
0.3310
0.3073
0.2645
20.59%
-
Average
0.6077
0.0920
0.3551
0.3308
0.2933
45.49%
-
Qwen-VL-Chat
No corrupt
0.4389
0.0806
0.2562
0.0000
0.2118
38.90%
100.00%
TABLE II : Performance comparison on Stanford Dogs Caption under different corruption strategies.
Model
Strategy
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Dec. Rate
BFG(ours)
No corrupt
0.0521
0.1063
0.3005
0.3055
0.8914
-
Replace
0.0516
0.1051
0.2993
0.3021
0.8915
-
Drop
0.0636
0.1064
0.3023
0.3046
0.8926
-
Repeat
0.0528
0.1062
0.2977
0.3053
0.8918
-
Average
0.0550
0.1060
0.2999
0.3044
0.8918
-
Qwen-VL-Chat
No corrupt
0.0043
0.0196
0.1812
0.2011
0.8634
100.00%
TABLE III : Performance comparison on Flickr30k under different corruption strategies.
Fig. 5 : CIDEr score comparison and retention rate heatmap on the Stanford Dogs Caption dataset under different corruption strategies: (a) CIDEr score comparison across models and corruption strategies; (b) CIDEr retention rate heatmap showing the percentage of baseline performance maintained.
Fig. 6 : CIDEr score comparison and retention rate heatmap on the Flickr30k dataset under different corruption strategies: (a) CIDEr score comparison across models and corruption strategies; (b) CIDEr retention rate heatmap showing the percentage of baseline performance maintained.
Fig. 7 : Robustness comparison of BFG and vision-language models on Flickr30k. Each axis represents one corruption condition. Top row is BERTScore-F1 and bottom row is METEOR. BFG covers the largest area across all conditions.
Fig. 8 : Visual comparison of model-generated descriptions under 2% byte replacement.
With byte masking
Without byte masking
Metric
Clean
Replace
Drop
Repeat
All corrupt
Clean
Replace
Drop
Repeat
All corrupt
CIDEr
0.6339
0.6240
0.6248
0.5479
0.5989
0.6473
0.6336
0.6347
0.5637
0.6107
BLEU-4
0.0996
0.0966
0.0970
0.0748
0.0895
0.1086
0.1028
0.1046
0.0720
0.0931
ROUGE-L
0.3651
0.3623
0.3621
0.3310
0.3518
0.3755
0.3680
0.3702
0.3262
0.3548
METEOR
0.3408
0.3375
0.3376
0.3073
0.3275
0.3506
0.3427
0.3453
0.3020
0.3300
Breed Acc.
57.63%
51.24%
52.50%
20.59%
41.44%
62.94%
55.02%
57.44%
20.18%
44.21%
TABLE IV : Ablation on byte-level corruption augmentation for Stanford Dogs Caption. Bold values indicate the smaller relative drop for each metric and corruption condition.
With byte masking
Without byte masking
Metric
Clean
Replace
Drop
Repeat
All corrupt
Clean
Replace
Drop
Repeat
All corrupt
CIDEr
0.0521
0.0537
0.0560
0.0548
0.0549
0.0498
0.0544
0.0543
0.0465
0.0517
BLEU-4
0.1063
0.1060
0.1056
0.1070
0.1062
0.1046
0.1030
0.1042
0.0924
0.0999
ROUGE-L
0.3005
0.3004
0.3021
0.2975
0.3000
0.2956
0.2950
0.2963
0.2860
0.2925
METEOR
0.3055
0.3034
0.3031
0.3078
0.3048
0.3019
0.3003
0.3007
0.2983
0.2998
BERTScore-F1
0.8914
0.8915
0.8926
0.8918
0.8920
0.8900
0.8897
0.8899
0.8868
0.8888
TABLE V : Ablation on byte-level corruption augmentation for Flickr30k. Bold values indicate the smaller relative drop for each metric and corruption condition.
Format
CIDEr
BLEU-4
ROUGE-L
METEOR
BERTScore-F1
Breed Acc.
JPEG
0.6339
0.0996
0.3651
0.3408
0.3058
57.63%
WebP
0.4077
0.0488
0.2913
0.2748
0.1944
0.44%
PNG
0.3381
0.0526
0.2934
0.2787
0.1964
0.44%
TABLE VI : Zero-shot evaluation on additional image formats.
Fig. 9 : Visual comparison of model-generated descriptions on Flickr30k different bitstream corruption.
Resource-constrained visual Internet of Things (IoT) systems, such as edge cameras, unmanned sensing platforms, industrial inspection nodes, and remote monitoring sensors, often need to transmit task-relevant visual evidence over low-rate wireless links to an edge/cloud service. Existing image communication methods usually compress or transmit complete global representations, leaving limited room to exploit receiver-side generative restoration. This paper proposes a semantic-aware generative image transmission framework for edge-assisted visual IoT. The image captured by an IoT visual sensor is encoded into a discrete token grid by a VQ encoder. At the IoT transmitter or nearby gateway, token recoverability, estimated from prediction entropy and local structure complexity, is fused with semantic importance obtained from instance segmentation and category-aware scoring. A spatial dispersal sampler then selects the tokens to be transmitted under a bitrate budget. The transmitter sends only the quantization indices of kept tokens and a binary mask map, while the edge/cloud receiver recovers masked tokens through MaskGIT with Halton sequence scheduling. Experiments on Kodak and VisDrone scenes under AWGN and Rayleigh channels show that the proposed method provides a flexible bitrate-quality tradeoff for narrowband visual IoT links. At 0.074 bpp, it uses 44.6% of the transmitted bits of the 0.167-bpp DeepJSCC/WITT reference while achieving 29.9 dB PSNR. A pseudo-GT downstream detection study on Kodak further shows that semantic-aware masking preserves task-relevant objects better than random masking at both 30% and 50% mask ratios.
Chenyang Zhang, Changwang Liu, Jinqi Zhu +4
School of Computer and Information Engineering, Tianjin Normal University, Tianjin, China · School of Information Science and Engineering, Linyi University, Linyi 276000, China
AI-generated image (AIGI) detectors achieve strong accuracy on clean benchmarks, but their performance drops sharply after images are propagated through real-world channels. We trace this fragility to what these detectors actually learn: they overfit to local artifacts left by generators in small spatial neighborhoods, which are easily destroyed by common propagation degradations such as JPEG compression and blur. Instead, we shift the discriminative cue from fragile local artifacts to more robust global structure. Building on this, we propose GlobalForge, a framework with two complementary modules. The Local Information Bottleneck (LIB) suppresses local components to block shortcut learning, while the Global Structural Reasoning (GSR) module forces every token to gather evidence from distant regions. Both modules are trained jointly under a contrastive structural loss based on degradation that keeps the resulting features stable under degradation. To support fine-grained robustness evaluation, we further introduce RealDeg-Bench, covering 7 common degradation operators and multi-step compound chains. GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by 5.89% over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations. Code is available at https://anonymous.4open.science/r/GlobalForge-BE0F/.
Manni Cui, Ruiqi Liu, Dianyuan Zou +8
1Huazhong University of Science and Technology · Institute of Automation, Chinese Academy of Sciences · 3Jilin University +1
Detecting AI-generated images across unseen architectures remains challenging, as existing models often overfit to generator-specific fingerprints and semantic content rather than learning universal forgery traces. We attribute this failure to feature entanglement: detectors learn these factors as a single entangled representation, where universal forgery traces are inextricably confounded with both generator-specific fingerprints and semantic content. Crucially, our spectral analysis reveals that this entanglement is avoidable: distinct generator-specific fingerprints (e.g., GAN stripes vs. Diffusion Model spots) occupy disjoint frequency subspaces and coexist as independent superpositions. Leveraging this physical orthogonality, we propose the Orthogonal Decomposition and Purification Network (ODP-Net) to structurally disentangle these factors. Specifically, ODP-Net employs (1) Instance-aware Orthogonal Decomposition to project features into mutually exclusive subspaces: universal forgery traces, generator-specific fingerprints, and semantic content; (2) Perturbation-based Purification to enforce semantic invariance via cross-sample feature injection; and (3) Manifold Alignment to bridge domain gaps. By explicitly decoupling universal forgery traces from generator-specific fingerprints and semantic content, ODP-Net achieves state-of-the-art performance on unseen architectures (e.g., Stable Diffusion 3), validating that structural disentanglement is key to generalization.
Zhiyuan Wang, Yanxiang Chen, Pengcheng Zhao +2
Hefei University of Technology · Key Laboratory of Knowledge Engineering with Big Data · Key Laboratory of Knowledge Engineering with Big Data (Hefei University of Technology), Ministry of Education; School of Computer Science and Information Engineering, Hefei University of Technology; and Intelligent Interconnected Systems Laboratory of Anhui Province (Hefei University of Technology) +3