cs.AIAug 21, 2026

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

Authors: Alexander Thomas, Hubert P. H. Shum, Darren Nellis, Manli Zhu, Phatpicha Yochum, William Bartle, Daniel Wrightson

Organizations: NCB Hazcheck Limited, UK · Durham University, UK

Abstract

The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on a commercial e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.

Figures & tables

Explore similar work

CardsList
  1. When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

    May 26, 2026Yige Li, Jun Sun, Wei Zhao +5Large Language Model SafetyStress Testing

  2. IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

    May 11, 2026Songlin Bai, Xintong Wang, Linlin Yu +12Safety BenchmarksLarge Language Model Evaluation

  3. Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

    May 7, 2026Haoxiang Wang, Da Yu, Huishuai ZhangLarge Language Model EvaluationPass-Rate Evaluation