cs.CVMay 10, 2025

Batch Augmentation with Unimodal Fine-tuning for Multimodal Fusion of Large Language Models

Authors: H M Dipu Kabir, Subrota Kumar Mondal, Mohammad Ali Moni

Organizations: AI and Cyber Futures Institute, Charles Sturt University, Australia · Rural Health Research Institute, Charles Sturt University, Australia · School of Computer Science and Engineering, Macau University of Science and Technology, Macao

Abstract

In this paper, we propose batch augmentation with unimodal fine-tuning for multimodal learning. We start with pre-trained unimodal models. We fine-tune the unimodal models with the application data. After that, we form a Multi-Layer Perceptron (MLP) head that takes information from unimodal models and provides output. Finally, we train the MLP layer and unimodal parts with batch augmentation. Depending on the data, some unimodal models can be replaced by hard-coded scripts or AI agents. The unimodal training can also follow batch augmentation when the data is augmentable. We write a multimodal batch augmentation dataloader script that implements the batch augmentation for the multimodal data. We investigate the proposed method on the FPU23 ultrasound and UPMC Food-101 multimodal datasets. The multimodal large language model (LLM) with the proposed training achieves the best average result among the investigated methods across both datasets. According to our literature search, the proposed method achieves state-of-the-art (SOTA) accuracy of 93.29% on the UPMC Food-101 dataset, while we apply the ViT-L/16 model for vision and the GPT-2 model for text. We share the scripts of the proposed method with traditional counterparts at the following repository: github.com/dipuk0506/multimodal

Explore similar work

CardsList
  1. Fusion Anything: A Generalized Multimodal Foundation Model

    Aug 16, 2026Huizi Cui, Zongbo Han, Chenggong Ding +6Multimodal FusionMultimodal Foundation Model

  2. Pareto LoRA: Mitigating Modality Imbalance in Unified Multimodal Models via Pareto-Optimal Gradient Integration

    Jun 15, 2026Xiwen Wei, Mark Nutter, Madhusudhanan Srinivasan +1Multimodal Continual Instruction TuningModalities

  3. PQFA: Parallel Quantum Feature Augmentation of Fused Representations for Multimodal Classification

    Jul 15, 2026Mingzhu Wang, Yun ShangMultimodal ClassificationHybrid Quantum-Classical Pipeline