NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework
Organizations: Harbin Engineering University, 145 Nantong Street, Nangang District, Harbin, P. R. China
Abstract
Ship-form design combines smooth geometric representation, local shape editing, and constraints on the resulting hull. We present the Natural-Language-to-Hull Framework (NL2Hull Framework), which formulates ship-form editing as a typed discrete decision problem and connects language decisions to numerical geometry. Its Constrained Free-Form Deformation Engine (CFFD Engine) represents hull waterlines with non-uniform rational B-splines (NURBS), applies free-form deformation (FFD) to their control points, reconstructs the hull, and checks geometric constraints. We construct the Ship Design Decision Dataset (SDD Dataset) with 134,558 cleaned records and evaluate compared models on its subset Ship Design Decision Benchmark (SDDBench), containing 5,000 records and 43,496 typed questions. We propose Chip, a constrained ship-design decision model for processing natural-language requests. Chip reaches 95.90% question accuracy and 99.32% FFD exact match, with a negative log-likelihood of 0.0951, an expected calibration error of 0.0032, and a Brier score of 0.0551. The NL2Hull Framework provides a reproducible interface for evaluating language-based ship-form decisions while identifying the geometry and continuous-control components that require further development. Our code and dataset is available at https://github.com/wenhuahuo/NL2Hull.
Figures & tables
| Hull | Surface RMSE | Waterline width RMSE | Fore-profile RMSE | Aft-profile RMSE |
| DTC | 0.01130 | 0.00302 | 0.00038 | 0.00121 |
| DTMB5415 | 0.00524 | 0.00227 | 0.00070 | 0.00072 |
| KCS | 0.00623 | 0.00312 | 0.00034 | 0.00082 |
| KVLCC2 | 0.00975 | 0.00140 | 0.00012 | 0.00092 |
| Containership | 0.00690 | 0.00202 | 0.00003 | 0.00038 |
| Firgate | 0.00658 | 0.00284 | 0.00023 | 0.00084 |
| Operation | Level | Volume ratio | Watertight | Valid volume |
| Bow outward | 5 | 1.037 | Yes | Yes |
| Bulb forward | 4 | 1.033 | Yes | Yes |
| Global breadth increase | 5 | 1.120 | Yes | Yes |
| Model | Question accuracy | FFD exact match | NLL | ECE | Brier |
| Small language models | |||||
| Qwen3 | 75.61 | 2.84 | 1.2088 | 0.1148 | 0.3803 |
| Gemma3 | 66.97 | 15.98 | 1.4328 | 0.0225 | 0.4471 |
| Llama3.2 | 63.94 | 2.16 | 1.4999 | 0.0723 | 0.5063 |
| Flagship language models | |||||
| Claude Opus 5 | 92.87 | 99.12 | 0.2385 | 0.0607 | 0.1066 |
| Model | Easy | Original | Hard | Brier | ECE |
| JEV | 100.0000 | 98.6100 | 72.0700 | 0.3584 | 0.0947 |
| Laya | 95.8300 | 72.2200 | 28.8300 | 0.8042 | 0.2465 |
| SemIf | 100.0000 | 98.6100 | 61.2600 | 0.4980 | 0.1122 |
| JevK5 | 100.0000 | 97.2200 | 73.8700 | 0.3662 | 0.0467 |
| Intern-Decision-0.8B | 97.9200 | 80.5600 | 52.2500 | 0.5295 | 0.0657 |
| Chip (Ours) | 100.0000 | 81.9444 | 36.9369 | 0.3489 | 0.1633 |
| Interaction type | Experiments | End-to-end success | Trajectory exact | Geometry valid | Constraints satisfied | Level accuracy |
| Single edit | 100 | 100.00% | 100.00% | 100.00% | 100.00% | 52.00% |
| Multi-region, single turn | 100 | 60.00% | 96.00% | 100.00% | 13.33% | 70.48% |
| Same-region, two turns | 100 | 98.00% | 100.00% | 100.00% | 95.35% | 81.50% |
| Model | Mean latency (ms) | 95th percentile (ms) |
| Small language models | ||
| Qwen3 | 25280.05 | 46745.65 |
| Gemma3 | 18407.59 | 43119.55 |
| Llama3.2 | 22589.78 | 44708.37 |
| Flagship language models | ||
| Claude Opus 5 | 7844.00 | 12531.86 |
| Variant | 32 controls | Deck knot | Bulb knots | KCS stem RMSE | KCS stern RMSE | Mean stem RMSE | Mean stern RMSE |
| 24 controls | – | – | – | 0.002326 | 0.001581 | 0.001551 | 0.001236 |
| 32 controls | ✓ | – | – | 0.001222 | 0.000818 | 0.000909 | 0.000785 |
| Deck knot | – | ✓ | – | 0.002187 | 0.001581 | 0.000994 | 0.001236 |
| Bulb knots | – | – | ✓ | 0.001684 | 0.001581 | 0.001396 | 0.001236 |
| Combined | ✓ | ✓ | ✓ | 0.000343 | 0.000818 | 0.000221 | 0.000785 |
| Model | Completed records | Question accuracy | FFD exact match | NLL | ECE | Brier score |
| Chip | 5,000 | 95.8962% | 99.32% | 0.0951 | 0.0032 | 0.0551 |
| Chip-2B | 5,000 | 95.9192% | 99.30% | 0.0948 | 0.0022 | 0.0548 |
| Chip-4B | 4,999 | 95.8794% | 99.24% | 0.0951 | 0.0025 | 0.0549 |
| Training share (%) | Training records | Question accuracy | FFD exact match | NLL | ECE | Brier |
| 0 | 0 | 80.82 | 72.24 | 0.6972 | 0.1995 | 0.3439 |
| 1 | 886 | 94.02 | 96.86 | 0.1729 | 0.0179 | 0.0850 |
| 2 | 1,772 | 94.57 | 97.88 | 0.1533 | 0.0110 | 0.0754 |
| 4 | 3,544 | 95.13 | 98.64 | 0.1331 | 0.0105 | 0.0682 |
| 8 | 7,088 | 95.39 | 98.82 | 0.1196 | 0.0097 | 0.0642 |
| 25 | 22,151 | 95.66 | 99.14 | 0.1040 | 0.0047 | 0.0584 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Model | Language backbone | Batch size | Accumulation |
| Chip | Qwen3.5-0.8B-Base | 2 | 4 |
| Chip-2B | Qwen3.5-2B-Base | 2 | 4 |
| Chip-4B | Qwen3.5-4B-Base | 1 | 8 |