Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
Abstract
Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.
Figures & tables
| Metric | Description | Application |
| Hamming ( ) | Number of substituted tokens (for strings of equal length) | : 1–5; token substitution attacks |
| Levenshtein ( ) | Min. number of operations (insertion, deletion, substitution) | : 1–10; typos, short insertions |
| Semantic ( ) | Cosine distance between embeddings | : 0.05–0.2; paraphrasing |
| ID | Threat (OWASP) | Description | Threats from Sec. 1.2.2 |
|---|---|---|---|
| LLM01 | Prompt Injection | Injection of malicious instructions that override the behavior of the model. | 1. Prompt injection, 14. Jailbreaking |
| LLM02 | Insecure Output Handling | Insufficient validation and sanitization of LLM outputs before use in other systems. | 9. Insecure output handling |
| LLM03 | Training Data Poisoning | Manipulation of the training data to compromise the model. | 2. Data poisoning, 3. Backdoor attacks |
| LLM04 | Model Denial of Service (Model DoS) | Attacks that lead to resource exhaustion and model unavailability. | 7. Denial of service (DoS) |
| LLM05 | Supply Chain Vulnerabilities | Use of vulnerable components: models, dependencies, data. | 13. Supply chain vulnerabilities |
| LLM06 | Sensitive Information Disclosure | Unintended disclosure of confidential data by the model. | 11. Sensitive information disclosure, 4. Model inversion, 5. Membership inference, 10. Data extraction |
| Threat | NIST control families | Example controls |
|---|---|---|
| 1. Prompt injection | SI, AC | SI-10, AC-6, AC-3 |
| 2. Data poisoning | SI, SR | SI-7, SR-3, SR-6 |
| 3. Backdoor attacks | SI, SR | SI-7, SR-4, SR-5 |
| 4. Model inversion | AC, PT | AC-3, AC-4, PT-3, PT-4 |
| 5. Membership inference | PT, SI, AU | PT-3, SI-4, AU-6 |
| 6. Adversarial examples | SI, AT | SI-4, SI-10, AT-2 |
| Metric | Invar. | Sens. | Comput. |
| High | High | Med./High | |
| High | High | High | |
| BLEU/ROUGE-like metrics | Medium | Low | High |
| Perplexity-based score | Low | Medium | High |
| Model | BI | CWB | CM | ASA |
| Gemma-3-27B-Instruct | 57.5 6.9 | 48.7 6.9 | 67.7 6.5 | 72.3 6.2 |
| Gemma-3-4B-Instruct | 66.7 6.5 | 55.5 6.9 | 67.4 6.5 | 73.8 6.1 |
| Llama-3.2-3B-Instruct | 60.1 6.8 | 47.5 6.9 | 47.8 6.9 | 58.2 6.8 |
| GPT-4 | 32.4 6.5 | 28.6 6.3 | 41.2 6.8 | 45.7 6.9 |
| Claude-3-Opus | 29.8 6.3 | 25.3 6.0 | 38.5 6.7 | 42.9 6.9 |
| Defense mechanism | ASR (% CI) |
| Perplexity Check | 58.4 6.8 |
| Instruction Filtering | 63.7 6.7 |
| Content Moderation | 67.5 6.5 |
| All defenses combined | 42.1 6.8 |
| Committee of models | |
| 3 Models | 39.4 6.8 |
| Method | Recall | REASR |
| Random baseline * | 0.14 | – |
| PEZ (baseline) | 0.105 | 0.052 |
| GBDA (baseline) | 0.116 | 0.056 |
| UAT (baseline) | 0.131 | 0.030 |
| GCG (best run) | 0.167 | 0.987 |
| Metric | Value |
| (mean std. dev.) | |
| Yes | |
| Fraction of preserved discrete outputs | 5% |
| Number of prompts | 20 |
| Number of perturbations per prompt | 5 |
| Characteristic | HF Pipeline | vLLM | Gain |
| Throughput (req./s) | 2.8 | 14.1 | 5.0 |
| Latency (p50), ms | 1420 | 310 | 4.6 |
| Latency (p99), ms | 3200 | 780 | 4.1 |
| GPU memory utilization, % | 62 | 91 | +29 pp |
| Max. parallel requests | 4 | 32 | 8.0 |
| Full evaluation time (1000 prompts) * , min | 47.2 | 9.8 | 4.8 |
| Attack type | Baseline ASR | MCP ASR | |
| Indirect injection | 31.2% | 47.8% | +16.6 pp |
| Tool response manipulation | 28.4% | 52.1% | +23.7 pp |
| Cross-server propagation | 19.7% | 61.3% | +41.6 pp |
| Average | 26.4% | 53.7% | +27.3 pp |
| Component | Latency, ms | CO share, % |
| Base inference (vLLM) | 310 | – |
| Perplexity filter | 42 | 13.5 |
| Content moderation | 28 | 9.0 |
| AttestMCP attestation | 0.08 | 0.03 |
| Metric computation | 15 | 4.8 |
| Logging and storage | 3 | 1.0 |
| Abbreviations | |
| AGM | Attack Generator Module |
| AI | Artificial Intelligence |
| AI RMF | NIST AI Risk Management Framework |
| API | Application Programming Interface |
| AR | Attack Robustness |
| ASA | Adaptive Search-Based Attack |