cs.CRSep 27, 2026

COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints

Authors: Hao Chen

Organizations: School of Cyber Security, Guangdong Polytechnic Normal University, Guangzhou, China

Abstract

When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probability calibration, dual-use false-positive control, and heterogeneous CPU-NPU routing under explicit latency SLOs. Pre-ingestion guardrails must screen prompts prior to target-LLM prefill with low false alarms on benign compliance inquiries; however, shallow classifiers are brittle to phrasing shifts, hidden-state probes require coupling to a target LLM, and generative guards incur high decoding latency and dual-use false positives. We present COGNIT-Guard, coupling a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to an NPU-resident 322M bidirectional direct-decision model (Laya-322M) under an asymmetric false-positive penalty. On the clean unseen DUCS-Bench test split (N=607N=607), COGNIT-Guard achieves 98.85% accuracy (McNemar p=1.19×10−4p = 1.19 \times 10^{-4} vs. ML), reduces benign FPR to 0.42% (1/2381/238; Fisher's exact p=8.23×10−4p = 8.23 \times 10^{-4} vs. ML), and attains 1.12% ECE and 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference runs in 21.77 ms mean latency (45.90 QPS), while the live serial CPU-NPU cascade (θdeploy∗=0.70θ^*_{\mathrm{deploy}}=0.70) achieves 41.63 ms mean latency (P50: 39.47 ms, 99.23% accuracy, 0.00% FPR). Evaluation on SafetyBench-ZH (N=2,100N=2,100) and comparison against a bi-encoder direct-decision baseline (CLM-8B) disentangle in-domain gains, OOD alignment tax (60.33% →\to 56.81% on Laya; 55.10% on domain CLM-8B), and experience replay recovery, restoring OOD accuracy to 64.10%-65.05% and reaching 99.67%-99.84% in-domain accuracy with 0.00%-0.42% FPR.

Figures & tables

Explore similar work

CardsList
  1. Robust and Efficient Guardrails with Latent Reasoning

    May 27, 2026Siddharth Sai, Xiaofei Wen, Muhao ChenLarge Language Model SafetyGuardrail

  2. kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

    Jul 2, 2026Mahmoud Abdelfattah, Hamid Nasiri, Peter GarraghanLarge Language Model SafetyAdversarial Prompts