COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints
Organizations: School of Cyber Security, Guangdong Polytechnic Normal University, Guangzhou, China
Abstract
When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probability calibration, dual-use false-positive control, and heterogeneous CPU-NPU routing under explicit latency SLOs. Pre-ingestion guardrails must screen prompts prior to target-LLM prefill with low false alarms on benign compliance inquiries; however, shallow classifiers are brittle to phrasing shifts, hidden-state probes require coupling to a target LLM, and generative guards incur high decoding latency and dual-use false positives. We present COGNIT-Guard, coupling a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to an NPU-resident 322M bidirectional direct-decision model (Laya-322M) under an asymmetric false-positive penalty. On the clean unseen DUCS-Bench test split (), COGNIT-Guard achieves 98.85% accuracy (McNemar vs. ML), reduces benign FPR to 0.42% (; Fisher's exact vs. ML), and attains 1.12% ECE and 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference runs in 21.77 ms mean latency (45.90 QPS), while the live serial CPU-NPU cascade () achieves 41.63 ms mean latency (P50: 39.47 ms, 99.23% accuracy, 0.00% FPR). Evaluation on SafetyBench-ZH () and comparison against a bi-encoder direct-decision baseline (CLM-8B) disentangle in-domain gains, OOD alignment tax (60.33% 56.81% on Laya; 55.10% on domain CLM-8B), and experience replay recovery, restoring OOD accuracy to 64.10%-65.05% and reaching 99.67%-99.84% in-domain accuracy with 0.00%-0.42% FPR.
Figures & tables
| System / Guardrail Family | Moderation Paradigm | Standalone Pre-Ingestion | Target-LLM Decoupled | Mean Latency | Dual-Use FP Control / Calibration | Heterogeneous Routing / Hardware |
|---|---|---|---|---|---|---|
| Llama-Guard 3 [ 1 ] / WildGuard [ 2 ] / ShieldGemma [ 3 ] | Autoregressive Guard (7B–70B) | Yes | Yes | – ms | Uncalibrated Token Logits ( – XSTest FPR) | No (Single GPU Tier) |
| ShieldHead [ 4 ] / SingProbe [ 5 ] / Khatri et al. [ 6 ] | Intrinsic Hidden-State Probe | No (Requires LLM Prefill) | No (Coupled to Base LLM) | – ms (Prefill) | Uncalibrated on Dual-Use Compliance | No (Co-located on Target GPU) |
| GLiNER Guard [ 7 ] / SingGuard [ 8 ] / FlexGuard [ 9 ] | Single-Pass Encoder / Risk Scorer | Yes | Yes | – ms | Continuous Score ( XSTest FPR) | No (Single Accelerator Tier) |
| CLM-8B (Bi-Encoder Direct Baseline) [ 10 ] | Disaggregated Bi-Encoder Head | Yes | Yes | ms ( -Str.) / ms (Batch) / ms (Arena) § | (Zero-Shot) 0.00% (Rehearsal, ECE) | HBM VectorArena (Ascend 910C) |
| Traditional ML Baseline | TF-IDF + XGBoost ( M) | Yes | Yes | ms (Diag.) / ms (Fast CPU) † | Poorly Calibrated ( ECE, DUCS FPR) | CPU Only (Xeon CPU) |
| COGNIT-Guard (Ours) | Calibrated Direct Cascade (322M) | Yes | Yes | 14.27 ms (Offline) / 21.77 ms (NPU) / 41.63 ms (Live E2E) ‡ | 0.00%–0.42% DUCS FPR (1.12% ECE) | Calibrated CPU–NPU Cascade (Ascend 910C) |
| Evaluation Metric | Traditional ML | Zero-Shot DFM | COGNIT-Guard (Ours) |
|---|---|---|---|
| Strict Clean Unseen Partition ( : Benign, Violations) | |||
| Accuracy | 95.39% (579/607) | 76.61% (465/607) | 98.85% (600/607) |
| [95% Wilson CI] | [93.38%, 96.81%] | [73.08%, 79.82%] | [97.64%, 99.44%] |
| Paired McNemar (Acc.) | Baseline | (***) | |
| False Positive Rate (FPR) | 5.88% (14/238) | 35.71% (85/238) | 0.42% (1/238) |
| [95% Wilson CI] | [3.52%, 9.61%] | [29.88%, 41.98%] | [0.07%, 2.34%] |
| Configuration / Execution Stage | NPU Ratio | System Acc. | FPR | Mean Latency | P50 / P95 Lat. |
|---|---|---|---|---|---|
| Part A: Offline Model-Compute Sweep ( , , Pre-Extracted Cached Features) | |||||
| Pure NPU Path ( ) | 100.0% | 98.90% | 0.42% | 21.76 ms | 21.53 / 23.58 ms |
| Cascade ( ) | 61.3% | 98.58% | 0.00% | 17.40 ms | 17.10 / 22.71 ms |
| Cascade ( ) | 33.7% | 98.58% | 0.00% | 14.27 ms | 10.50 / 21.83 ms |
| Cascade ( ) | 8.3% | 97.32% | 2.52% | 11.47 ms | 10.50 / 21.49 ms |
| Mode C: Cached-Feature Tree ( ) | 0.0% | 97.17% | 2.52% | 10.50 ms | 10.50 / 10.50 ms |
| Execution Strategy (Ascend 910C) | Mean Latency | P95 Latency | Cold-Start / Alloc Overhead | Speedup / Throughput |
|---|---|---|---|---|
| Laya-322M: Cold-Start Unpinned Execution | 48.20 ms | 62.50 ms | 14.80 ms | 1.00 (20.7 QPS) |
| Laya-322M: Pre-Allocated Memory Buffer | 32.40 ms | 41.10 ms | 6.20 ms | 1.49 (30.9 QPS) |
| Laya-322M: Warm On-Chip HBM Residency (Ours) | 21.77 ms | 23.44 ms | 0.00 ms | 2.21 (45.9 QPS) |
| CLM-8B Baseline: Uncached Single-Stream ( ) | 45.78 ms | 48.51 ms | 0.00 ms | 21.8 QPS ( ) |
| CLM-8B Baseline: Batched ( ) / HBM VectorArena | 6.64 / 0.60 ms | 7.12 / 0.89 ms | 0.00 ms | 150.6 / QPS |
| System / Evaluated Model | SafetyBench Acc. | SafetyBench FPR | DUCS FPR (ZH) | XSTest FPR (EN) | Mean Latency |
|---|---|---|---|---|---|
| GPT-4 (Zero-Shot) [ 21 ] | 89.10% | – | 18.40% (46/250) | ms (API) | |
| Llama-2-70B-Chat [ 21 ] | 68.70% | – | 41.60% (104/250) | ms (GPU) | |
| WildGuard (7B Decoder) [ 2 ] | – | – | – | 7.20% (18/250) | ms (GPU) |
| SingGuard (Lightweight) [ 8 ] | – | – | – | 4.80% (12/250) | ms (GPU) |
| Traditional ML (Fast CPU Mode) | 53.29% (1119/2100) | 37.98% (534/1406) | 5.88% (14/238) | – | 22.79 ms (CPU) |
| Zero-Shot Cross-Enc. DFM (Laya-322M) | 60.33% (1267/2100) | 24.68% (347/1406) | 35.71% (85/238) | – | 44.85 ms (NPU) |