cs.CVSep 30, 2026

Comparative study of adapting pre-trained models for driving behavior video captioning

Authors: Sayak Mallick, Philipp Geiger, Augustin Kelava

Organizations: University of Tübingen · Bosch Center for Artificial Intelligence

Abstract

This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.

Figures & tables

Explore similar work

CardsList
  1. DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving

    May 22, 2026Hao Vo, Khoa Vo, Phu Loc Nguyen +10Autonomous DrivingSpatio-Temporal Reasoning

  2. Less Language, More Latents: Annotation-Efficient VLAs for Driving

    Sep 23, 2026Alexey Zakharov, Kemal Oksuz, Puneet K. DokaniaDiffusion-Based Vision-Language-ActionsLatent Action Models