cs.LGOct 6, 2026

Multi-Label Perceptual Bug Detection in Video Games using Deep Learning on Gameplay Footage

Authors: Nahian Rifaat, Felix Morosov, Loutfouz Zaman

Organizations: Ontario Tech University Oshawa, Ontario, Canada · Universität Osnabrück Osnabrück, Niedersachsen, Germany

Abstract

Traditional approaches for automated bug detection in video games, such as manual testing, can be beneficial for the improvement of quality assurance, but they can be expensive and time-consuming. The scarce number of tools available to detect multiple perceptual bugs in the same video frame introduces detection challenges for automated bug detection tools in real-world scenarios. We propose a deep learning model for multi-label perceptual bug detection and compare it against video classification models such as Inflated 3D ConvNet and 3D ResNet. Our proposed model, ResNet-BiLSTM, achieved an F1 score of 85.78% on the benchmark dataset. Our results demonstrated that temporal dependency modelling is beneficial for accurate video-based bug detection. We believe this work with multi-label perceptual bug detection on gameplay videos will help save resources spent on manual testing workloads in video games. Furthermore, we introduce a new dataset with multi-label perceptual bugs in this work. The dataset contains 77,969 video clips across different genres of games with approximately 1.2 million frames, containing combinations from 5 classes of bugs in the same video frame.

Figures & tables

Explore similar work

May 20, 2026cs.CV

TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos

Vision-language models (VLMs) are increasingly being explored for video game quality assurance, especially gameplay glitch detection. Most existing evaluations, however, treat glitches as static visual anomalies, asking models to detect failures from a single frame. We argue that this framing misses a key distinction: some glitches are spatial and visible in an isolated frame, whereas others are temporal and become evident only through changes across ordered frames. A preliminary study confirms this gap, showing that temporal glitches are substantially harder for VLMs to detect than spatial ones. To enable systematic evaluation of this underexplored setting, we introduce TempGlitch, a controlled gameplay video benchmark for temporal glitch detection. TempGlitch covers five temporal glitch types with balanced per-category samples, together with paired glitch-free videos that enable reliable binary evaluation. We evaluate 12 proprietary and open-weight VLMs across multiple frame-sampling settings. Our results show that current VLMs remain near chance on TempGlitch, often collapsing into either overly conservative behavior that misses most glitches or overly sensitive behavior that flags clean videos as glitchy. Moreover, denser frame sampling and larger model size do not reliably resolve these failures. TempGlitch provides a focused testbed for temporal reasoning, robust gameplay understanding, and automated glitch detection with VLMs. Code and data are available at the project website.
Apr 13, 2026cs.CV

RefGlitch-Bench: A Benchmark for Reference-based Gameplay Glitch Detection with Vision-Language Models

Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot keep pace with the growing test surface of modern game development. Prior automation efforts, particularly those using vision-language models (VLMs), largely operate on isolated frames without sufficient context to judge whether a glitch is present. We introduce RefGlitch-Bench, a benchmark for reference-based video game glitch detection with VLMs. The key idea is to formulate glitch detection as an explicit within-video comparison problem: given a test frame, a reference frame provides a visual baseline that helps the model distinguish true glitches from benign visual variation. RefGlitch-Bench includes a controlled synthetic dataset with five injected glitch types and manually annotated reference/test frame pairs, enabling an oracle-reference evaluation that isolates the potential benefit of reference guidance. We further establish four initial baselines for automatically selecting references from earlier frames in the same video, with LastCleanFrame performing best and transferring across VLMs. Finally, we evaluate automatic reference guidance on real-world gameplay data, where it improves frame-level glitch detection beyond the controlled setting while revealing reference reliability and error propagation as key challenges. Code and data are available at: https://github.com/PipiZong/RefGlitch-Bench.git.
Feb 16, 2026cs.CV

Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories

Reliable video understanding requires high-quality video datasets that can provide both precise semantic labels and temporally consistent annotations. Detecting annotation errors in densely labeled videos is challenging because errors may arise from semantic mislabeling, where labels disagree with visual content, or temporal disordering, where otherwise plausible labels violate procedural progression. Training dynamics have been used to identify mislabeled training examples primarily for static samples. We investigate checkpoint loss dynamics for out-of-sample auditing of temporally annotated videos. We compute Cumulative Sample Loss (CSL) as the mean annotation-conditioned loss of an audit frame across checkpoints trained on a disjoint reference set. CSL acts as a dynamic fingerprint and captures the persistent disagreement between its annotation and learned visual-temporal structure. High-CSL frames are then flagged as likely candidates for potential annotation errors, including semantic mislabeling or temporal disordering. Experiments on EgoPER and Cholec80 show that CSL substantially outperforms final-checkpoint loss and achieves up to a 4.2-point AUC improvement over prior baselines on EgoPER and 92.0/78.5 AUC for mislabeling/disordering on Cholec80. These results demonstrate checkpoint loss dynamics as an effective diagnostic for temporal annotation auditing.