Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
Organizations: PythaLab, Yıldız Technical University Istanbul, Turkey
Abstract
An acceptance protocol is developed for sensor-coordinate and polarity binding in mechatronic commissioning. Candidate generation is separated from release authority. Requirements unsupported by a deterministic parser are routed to a frozen local language model with four billion parameters. Plans are released only when both facts can be derived by an external gate under a sealed grammar. One canonical answer is requested from a gold-standard user when eligible. The protocol was evaluated once under a criterion fixed before benchmark construction, on 144 tasks written by isolated agent contexts without access to the gate, grammar, or experimental plan. Three contributions are established. First, candidate generation and release decisions were measured separately. Fabricated ready plans were committed on 21 of 22 routed unanswerable tasks, and all were rejected. The same 83 releases were reproduced without model calls. Second, no false release was observed among 83 releases. A one-sided 95% Clopper-Pearson upper bound of 0.0354 was obtained as a diagnostic under an independent-and-identically-distributed assumption, below the sealed 5% threshold. However, one false release was subsequently recorded among 146 releases outside the benchmark at seed 0. Third, protection against incorrect user answers was characterized. Both facts were bound from the original text on 13 of 96 answerable tasks. Incorrect answers were released in 169 of 431 pairings on the remaining tasks, including failures involving coordinate exclusion. A deployable questioning policy was not tested because eligibility was determined from the answer key. Gate sensitivity and real user behavior were not measured.
Figures & tables
| (a) Compositions by physical domain | |
|---|---|
| Physical domain | Compositions |
| Motion, actuation, and power transmission | Belt gantry tension, flywheel centrifugal brake, SMA wire tensioner, magnetic gear coupling, ball screw preloaded stage |
| Fluid power and hydrostatic bearings | Pump line accumulator, hydrostatic bearing pad, servo valve spool chamber |
| Thermal, process, and vacuum systems | Thermoelectric load plate, battery pack cold plate, recycle loop reactor, pressurised headbox, boiler drum natural circulation, vacuum chamber outgassing pumpdown |
| Vehicle, marine, flight, and hydropower systems | Wheel slip brake line, ship rudder yaw axis, short period pitch axis, penstock turbine governor |
| Microelectromechanical systems (MEMS) and precision transducers | Coriolis rate gyroscope, comb drive electrostatic actuator, inkjet piezo chamber |
| Component | Recorded value |
|---|---|
| Model | Frozen Qwen3, four billion parameters, 4-bit quantized weights |
| Local model server | Ollama 0.32.15 |
| Execution environment | Python 3.12.3 |
| Numerical libraries | NumPy 2.2.6, SciPy 1.14.0, and SymPy 1.13.3 |
| Generation settings | Temperature 0.6, nucleus sampling threshold (top-p) 0.95, candidate-token limit (top-k) 20 |
| Context / output limit | 12288 / 4096 tokens, reasoning mode enabled |
| Question | Experiment and arm | Table and figure |
|---|---|---|
| Question 1. Can a valid candidate be generated by the frozen model, and can its subsequent commitment as a ready plan be made correctly? | Main experiment, seed 17, second model | Table 5 , Figure 4 |
| Question 2. Is the false-release rate among released plans bounded under the sealed acceptance criterion? | Main experiment, two ungated arms | Table 4 , Figure 3 |
| Question 3. What shares of the delivered plans can be attributed to the model and the deterministic components? | Main experiment, model-free replay | Table 4 , Figure 5 |
| Question 4. Where is protection by the acceptance layer lost when an incorrect fact is supplied by the user, and how is that textual trust boundary defined? | Main experiment: wrong-answer and specificity sweeps. Earlier development experiment: selected strata | Table 7 , Figure 6 |
| Arm | Delivered | False | Correct | Upper bound | Lower bound | Model calls | Asked |
|---|---|---|---|---|---|---|---|
| Two-turn protocol, main experiment, seed 0 | — | ||||||
| Two-turn protocol, seed 17 | — | — | |||||
| Model-free replay | — | — | |||||
| Ungated shell, separate run | — | — | |||||
| In-run ungated terminal, seed 0 | — | — | |||||
| In-run ungated terminal, seed 17 | — | — |
| Qwen3, seed 0 | Qwen3, seed 17 | Second model | |
| 25 routed answerable tasks | |||
| Ready plan committed | 25 | 23 | 15 |
| Plan equal to gold | 24 | 22 | 11 |
| Validity among committed plans | 96.0% | — | 73.3% |
| Lower bound for equality to gold | 0.824 | 0.810 | — |
| Episode without a plan | 0 | 2 | 10 |
| Item | Count | Share or composition |
| (a) Turn-1 routing | ||
| Tasks executed at turn 1 | ||
| Deterministic branch | ||
| Model-calling branch | ||
| Answerable | ||
| Contradictory gold status, unanswerable | ||
| Regime | Measurement | Population | Upper bound |
| Rejected answer | 0 of 65 pairings | 13 protected tasks | |
| Trusted answer | 64 of 66 (97.0 %) | Protocol, silent substratum | — |
| 64 of 74 (86.5 %) | Gate, silent substratum | — | |
| 64 of 80 (80.0 %) | Protocol, all cell S | — | |
| Incorrect pair released | 4 of 290 pairings | Gate, cell C refusing | |
| 4 of 226 pairings | Protocol, cell C refusing |
| (a) Post-seal campaign, 252 tasks | ||
| Measurement | Seed 0 | Seed 17 (sensitivity) |
| Released plans | 146 | 143 |
| False releases | 1 | 0 |
| Upper bound | ||
| Lower bound | — | |
| Raw label, both tiers | supported | — |
| Approach | Location of authority | Basis of the decision |
|---|---|---|
| Simplex and runtime assurance ( Seto et al., 1998a ; Sha, 2001a ; Bak et al., 2009a ) | Verified decision module | Conditions on physical dynamics |
| Shielding ( Alshiekh et al., 2018a ) | Action-constraining shield | Formal shield synthesis |
| Candidate generator with an external critic ( Kambhampati et al., 2024a ) | External sound critic | Verification by the critic |
| Self-verification ( Stechly et al., 2025a ) | Candidate-generating model | Model-side checking of its own output |
| Present protocol | External deterministic entailment gate | Derivability of both facts from the requirement and user fact under grammar V1 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Arm | Design | Population | Measurement | Location |
|---|---|---|---|---|
| 1 | Seed-17 sensitivity run | Protected benchmark, seed 17 | Seed sensitivity of the sealed result | § 3.2 |
| 2 | Ungated shell, separate run | Protected benchmark, separate run | Ungated delivery and false release | § 3.2 |
| 3 | Specificity sweep | 4645 gate pairings | Ready status on mismatched requirements | § 3.4 |
| 4 | Wrong-fact instrument | Selected stratum from the earlier development experiment | Rejection of a text-contradicting fact | § 3.4 |
| 4b | Wrong-polarity instrument | Selected complementary stratum | Acceptance of a fact left open by text | § 3.4 |
| 5 | Model-free replay | Protected benchmark, deterministic | Measured model contribution | § 4 |
| Label | Composition / physical system | Role |
|---|---|---|
| T-1 | Magnetic gear coupling, a noncontact gear assembly | Only candidate-generation miss at seed 0. Proposed input speed , gold pole angle |
| T-2 | Thermoelectric load plate with a Peltier module | Seed-17 miss |
| T-3 | MEMS comb drive electrostatic actuator | Episode without a plan at seed 17 |
| T-4 | Wheel slip brake line in a vehicle | Release despite explicit exclusion |
| T-5 | Penstock and turbine governor | Release despite explicit exclusion |
| T-6 | Pressurised paper-machine headbox | Release without a question |
| Prediction | Measurement | Verdict |
| Arm 6, in-run ungated terminal | ||
| At least 1 false delivery and | 6 false of 13 deliveries over 144 turn-1 records, | Met |
| At least 6 false deliveries | 6 of 13 deliveries | Met at the threshold |
| At most 9 correct deliveries | 7 of 96 answerable tasks | Met |
| Arm 7, wrong-answer sweep | ||
| P1 0 releases in cell C refusing | 4 of 290 gate, 4 of 226 protocol | Falsified |
| Field | Sealed primary model | Second model, development-labeled |
| Model | Qwen3 with four billion parameters and 4-bit quantized weights | 4-bit quantized SmolLM3 model in the three-billion-parameter (3B) class, distributed through Hugging Face |
| Model fingerprint | Fingerprint retained in the sealed record for matching weights to the archive | Distinct fingerprint retained in the sealed record |
| Local runtime version | 0.32.15 | 0.32.15 |
| Sampling temperature | 0.6 | 0.6 |
| Nucleus sampling threshold | 0.95 | 0.95 |
| Candidate-token limit | 20 | 20 |
| Quantity | Recorded value | Displayed form | Source |
| Clopper–Pearson bounds | |||
| Sealed final look | |||
| Arm 6, seed 0 | |||
| Arm 6, seed 0 | |||
| Arm 6, seed 17 | |||
| Arm 6, seed 17 | |||
| Component | Recorded value | Record date |
|---|---|---|
| Processor | Intel Core i7-12700H, 20 logical central processing units (CPUs) | 12 September 2026 |
| System memory | 31.2 GiB total reported to the server | 1 September 2026 |
| Graphics processing unit (GPU) | NVIDIA GeForce RTX 3060 Laptop GPU | 13 September 2026 |
| GPU memory / driver | 6144 MiB / 591.74 | 13 September 2026 |
| Platform | Linux 6.18.33.2, Windows Subsystem for Linux 2 (WSL2), x86_64, glibc 2.39 | 12 September 2026 |
| Research line | Architecture | Evidence setting | Present addition |
|---|---|---|---|
| Simplex and runtime assurance ( Seto et al., 1998a ; Sha, 2001a ; Bak et al., 2009a ) | Untrusted component behind a verified decision module | Control systems without a language-model component | Language-model component and measured false-release upper bound |
| Safety for learned components ( Phan et al., 2020a ; Alshiekh et al., 2018a ) | Neural Simplex: transfer to a safe controller. Shielding: action restriction | Runtime assurance and formal shield synthesis | Entailment from textual requirements, without a physical safety result |
| Candidate generator and external verifier ( Kambhampati et al., 2024a ; Stechly et al., 2025a ) | Model-generated candidates checked by an external sound critic | Reasoning and planning tasks | Sealed criterion, blind benchmark, false-release upper bound |
| Language model and classical planner ( Liu et al., 2023a ) | Model-generated PDDL description solved by a classical planner | Plan generation and planning benchmarks | Release conditional on candidate derivability from text |
| Rails and control planes ( Rebedea et al., 2023a ; Reddy et al., 2026a ; Madatha, 2026a ) | Programmable checks, with some NeMo rails based on LLMs | Policy and safety checks | Deterministic entailment and a measured boundary for trust in user facts |
| Abstention ( Rajpurkar et al., 2018a ; Kirichenko et al., 2025a ; Wen et al., 2025a ) | Model-side abstention | Question answering, language-model benchmarks, and a method survey | Candidate generation and commitment measured on the same routed set |
| Symbol | Meaning |
|---|---|
| Tasks, texts, and sets | |
| A task | |
| Protected task set | |
| Natural-language requirement of task | |
| Sensor coordinates named in the task’s sensor cards | |
| Text evaluated by the gate | |