Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
Organizations: NCB Hazcheck Limited, UK · Durham University, UK
Abstract
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on a commercial e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.
Figures & tables
| Section | Name | Source | Question Type | Questions |
|---|---|---|---|---|
| 1 | MCQ Q&A | E-learning | Multiple choice (closed book) | 520 |
| 2 | Open-Ended Q&A | E-learning (adapted) | Free text (LLM-as-judge) | 360 |
| 3 | DGL Lookup | DGL | Structured lookup | 485 |
| 4 | Regulatory Recall | E-learning (derived) | IMDG section identification | 313 |
| Total | 1,678 |
| Model | Provider | Size | Thinking | S1 (%) | S2 (%) | S3 (%) | S4 (%) |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 4.6 | Anthropic | Small | None | 57.9 | 60.4 | 63.7 | 10.6 |
| Claude Sonnet 4.6 | Anthropic | Small | Low | 59.6 | 57.8 | 64.3 | 9.0 |
| Claude Sonnet 4.6 | Anthropic | Small | Medium | 59.9 | 60.8 | 64.0 | 11.4 |
| Claude Sonnet 4.6 | Anthropic | Small | High | 67.6 | 65.3 | 73.6 | 7.7 |
| Claude Sonnet 4.6 | Anthropic | Small | Max | 75.5 | 62.1 | 74.2 | 7.5 |
| Claude Opus 4.7 | Anthropic | Large | None | 65.8 | 68.2 | 77.5 | 17.4 |
| Model | Thinking | C/FF (%) | P/CH (%) | SL (%) | SO (%) |
|---|---|---|---|---|---|
| Claude Sonnet 4.6 | None | 67.8 | 69.7 | 52.2 | 54.8 |
| Claude Sonnet 4.6 | Low | 71.9 | 72.0 | 52.7 | 55.5 |
| Claude Sonnet 4.6 | Medium | 72.8 | 71.0 | 53.5 | 58.3 |
| Claude Sonnet 4.6 | High | 81.1 | 82.2 | 64.3 | 65.8 |
| Claude Sonnet 4.6 | Max | 84.4 | 82.3 | 70.8 | 73.5 |
| Claude Opus 4.7 | None | 77.4 | 80.3 | 61.2 | 62.5 |
| Model | Thinking | Cost / 1,000 Q ($) | Time / 1,000 Q |
|---|---|---|---|
| Claude Sonnet 4.6 | None | 0.70 | 42 min |
| Claude Sonnet 4.6 | Low | 0.61 | 40 min |
| Claude Sonnet 4.6 | Medium | 1.69 | 1 h 1 min |
| Claude Sonnet 4.6 | High | 6.43 | 2 h 21 min |
| Claude Sonnet 4.6 | Max | 15.91 | 4 h 57 min |
| Claude Opus 4.7 | None | 1.34 | 54 min |
| Section | Model | No Search (%) | Search (%) | Change (pp) | Cost Multiplier |
|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.6 (None) | 57.9 | 51.0 | -6.9 | 194× |
| 1 | Gemini 3.1 Flash-Lite (Minimal) | 65.9 | 72.2 | +6.3 | 7× |
| 1 | GPT-5.4 Mini (None) | 53.6 | 53.2 | -0.4 | 55× |
| 3 | Claude Sonnet 4.6 (None) | 63.7 | 69.3 | +5.6 | 538× |
| 3 | Gemini 3.1 Flash-Lite (Minimal) | 67.9 | 87.8 | +19.9 | ~1× |
| 3 | GPT-5.4 Mini (None) | 32.2 | 81.9 | +49.6 | 45× |
| Code | Title | Score | Weight | Relevant Roles |
|---|---|---|---|---|
| E51 | General Provisions | 4 | 0.027 | All |
| E52 | Classification - Substances & PSN | 6 | 0.041 | P/CH; SL |
| E53 | Packing Provisions | 6 | 0.041 | P/CH; SL |
| E54 | Consignment - Marks & Documents | 6 | 0.041 | Standard |
| E55 | Transport Operations - Stowage & Segregation | 6 | 0.041 | Standard |
| E55b | Transport Operations - Container Safety & Emergency | 6 | 0.041 | SL |
| Field | DGL Column | Description | Expert Score | Weight |
|---|---|---|---|---|
| UN Number | 1 | Four-digit UN identifier assigned to a dangerous good | - | 0.250 |
| PSN | 2 | Proper Shipping Name - official transport name | - | 0.250 |
| Primary Hazard Class | 3 | Hazard class, division, and compatibility group | 10 | 0.128 |
| Subsidiary Classes | 4 | Class numbers of any subsidiary hazards | 5 | 0.064 |
| Packing Group | 5 | I (most dangerous) to III (least dangerous) | 5 | 0.064 |
| Emergency Schedule | 15 | Emergency response codes for fire or spillage | 3 | 0.038 |