LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Figures & tables
Figure 1 : Why KaliBench Matters? KaliBench enables fine-grained, verifiable evaluation and runtime-free reward training for translating analyst intent into executable Kali commands.
Figure 2 : KaliBench end-to-end pipeline for NL-to-CLI evaluation on Kali Linux. The benchmark is constructed via LLM-based data generation, multi-round verification-regeneration, and representative-sample selection. Evaluation applies alias-aware parsing and fine-grained scoring over tools, arguments, and dimensional categories.
Figure 3 : KaliBench Structure and Representative Sample. KaliBench is a structured benchmark of Kali Linux tools organized along the security attack lifecycle. It comprises 5 security phases and 23 functional dimensions. The dataset sample includes natural-language queries and decomposed ground-truth commands. See Appendix D for a detailed view of examples of all 23 dimensions.
Figure 4 : KaliBench Training Protocol.
Model
Size
Exact Correct (%)
Tool Acc. (%)
Optional F1 (%)
Positional F1 (%)
Total Score (%)
(B)
U
R
H
U
R
H
U
R
H
U
R
H
U
R
H
Avg.
Compact Models ( ≤ 8B)
Foundation-Sec-Ins ⋆
8
14.7
19.1
56.2
66.5
92.5
86.8
33.5
37.4
75.6
53.0
60.0
81.3
51.0
63.3
81.2
65.2
Llama3.1-Ins
8
12.5
18.2
62.7
61.7
93.7
95.4
32.4
37.2
82.7
53.5
61.9
85.8
49.2
64.3
88.0
67.2
Llama-Primus-Base ⋆
8
15.8
19.3
60.3
67.3
90.7
94.1
36.6
39.2
80.8
56.3
61.9
84.0
53.4
63.9
86.3
67.9
Foundation-Sec-Rsn ⋆ †
8
13.3
18.6
66.6
67.0
92.8
95.9
34.8
39.1
85.5
53.0
57.9
86.2
51.6
63.3
89.2
68.0
Table 1 : Open-weight model query-to-command performance across three modes. Columns report Unrestricted (U), Restricted (R), and Hinted (H). † indicates thinking-enabled models, and ⋆ denotes cybersecurity-tuned models. Bold / underline mark the best and second-best within each size group. Models are sorted by Averaged Total Score. MoE parameters use Total/Active (e.g., 685/37B).
Model
Answered queries
EC (%)
EC (%) (all 5K)
GPT-5.6-Sol
4,988
61.83
61.68
Claude Opus 5
3,675
59.89
44.02
Codex (GPT-5.5)
4,633
55.77
51.68
Table 2 : Proprietary models and scaffolded inference in Unrestricted mode.
Figure 5 : Performance across security dimensions and modes. (a) Performance over 23 dimensions and 5 phases per mode. Box plots show distributions (Q1–Q3), with whiskers for min/max. (b,c) Total Score and Exact Correct (%) across modes; dotted lines indicate inter-mode gaps.
Query style
Ures.
Res.
Hinted
Paraphrase
+0.61
+0.22
-0.09
Informal
-0.55
-0.48
-0.28
Messy
-3.20
-1.40
-1.07
Table 3 : Average change in Total Score, in percentage points, relative to original queries across eight model configurations.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Overview of KaliBench samples. Representative examples spanning all five security phases and 23 evaluation dimensions are shown. Zoom in for finer details.
Model Name
Size
Inference
Hugging Face Repository
Llama3.1-Instruct
8B
local vLLM
meta-llama/Llama-3.1-8B-Instruct
Foundation-Sec-Instruct
8B
local vLLM
fdtn-ai/Foundation-Sec-8B
Foundation-Sec-Reasoning
8B
local vLLM
fdtn-ai/Foundation-Sec-8B-Reasoning
Llama-Primus-Base
8B
local vLLM
trend-cybertron/Llama-Primus-Base
Llama-Primus-Merged
8B
local vLLM
trendmicro-ailab/Llama-Primus-Merged
RedSage-Qwen3-8B-Ins
8B
local vLLM
RISys-Lab/RedSage-Qwen3-8B-Ins
Appendix
Table 4: Evaluated open-weight model configurations, parameter counts, inference platforms, and model sources. The inference platform used for each configuration is specified in the table.
Model
Size
Tool score Δ (pp)
Optional F1 Δ (pp)
Positional F1 Δ (pp)
Total score Δ (pp)
Exact Δ (pp)
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Compact Models ( ≤ 8B)
Foundation-Sec-Ins ⋆
8
26.00
-5.70
20.30
3.90
38.20
42.10
7.00
21.30
28.30
12.30
17.90
30.20
4.40
37.10
41.50
Llama3.1-Ins
8
32.00
1.70
33.70
4.80
45.50
50.30
8.40
23.90
32.30
15.10
23.70
38.80
5.70
44.50
50.20
Llama-Primus-Base ⋆
8
23.40
3.40
26.80
2.60
41.60
44.20
5.60
22.10
27.70
10.50
22.40
32.90
3.50
41.00
44.50
Foundation-Sec-Rsn ⋆ †
8
25.80
3.10
28.90
4.30
46.40
50.70
4.90
28.30
33.20
11.70
25.90
37.60
5.30
48.00
53.30
Appendix
Table 5 : Performance deltas across execution modes. We report differences between three execution modes: Unrestricted (U), Restricted (R), and Hinted (H). Δ1 denotes (R − U), Δ2 denotes (H − R), and Δ3 denotes (H − U), all measured in percentage points (pp). † indicates thinking-enabled models, and ⋆ denotes cybersecurity-tuned models. Bold / underline mark the largest and second-largest deltas within each size group. Model order follows Table 1 .
Figure 7 : Coverage of CLI options in KaliBench (Evaluation Subset). While common flags (e.g., --help , --verbose ) appear frequently, many functional tool-specific options also occur across the dataset.
Mode
RedSage-Ins
RedSage-K (GRPO)
ΔGRPO
RedSage-K (SFT+GRPO)
ΔSFT+GRPO
Unrestricted
56.36
67.40
+11.04
70.31
+13.95
Restricted
69.31
74.65
+5.34
77.93
+8.62
Hinted
89.53
88.77
-0.76
89.34
-0.19
Appendix
Table 6 : Effect of RLVR on RedSage-K across evaluation modes. Scores are Total Score percentages. Δ denotes the change from the original RedSage-Ins model in percentage points.
title
avg_total
avg_tool
avg_optional
avg_positional
avg_exact
count
score_sum
Best Reps
d2j-std-apk
1.000
1.000
1.000
1.000
1.000
72
5.000
rfdump
1.000
1.000
1.000
1.000
1.000
72
5.000
truecrypt2john
1.000
1.000
1.000
1.000
1.000
72
5.000
jsql
0.998
1.000
0.995
1.000
0.995
216
4.989
reconspider
0.998
1.000
1.000
0.993
0.993
288
4.984
Appendix
Table 7 : Best- and worst-representative tools aggregated from the scoring files of all models.
Model / Scaffold
Answered queries
Blocked / Missing
TA
Opt-F1
Pos-F1
TS
EC Answered
EC All 5K
GPT-5.6-Sol
4,988 (99.76%)
12 (0.24%)
92.54
87.42
71.51
83.83
61.83
61.68
Claude Opus 5
3,675 (73.50%)
1,325 (26.50%)
90.67
85.95
68.57
81.73
59.89
44.02
Codex (GPT-5.5, xhigh )
4,633 (92.66%)
367 (7.34%)
91.28
82.39
63.46
79.04
55.77
51.68
Appendix
Table 8 : Complete results for proprietary models and agentic scaffolds in the Unrestricted setting. We used the following abbreviations: Tool Accuracy (TA), Optional-Argument F1 (Opt-F1), Positional-Argument F1 (Pos-F1), Total Score (TS), and Exact Correct (EC). All metric scores are percentages. Bold shows the best.
Figure 8 : Model-wise Sensitivity to Messy Queries
We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier models from 7 providers on all 200 challenges of the NYU CTF Bench. Building on the D-CIPHER multi-agent framework, we extend it with multi-provider backend support, a custom Kali Linux environment with over 100 pre-installed penetration testing tools, and runtime tool-discovery agents. Through a controlled factorial study, we find that the Kali Linux environment yields a +9.5 percentage-point improvement over Ubuntu, while auto-prompting and category-specific tips often degrade performance in well-equipped environments. Among models, Claude 4.5 Opus achieves the highest solve rate (59%), followed by Gemini 3 Pro (52%), with Gemini 3 Flash offering the best cost-efficiency at $0.05 per solve. Asymmetric planner/executor model assignments provide no meaningful benefit while coherent same-model configurations consistently outperform mixed-tier pairings. Our results indicate that environment tooling and model selection emerge as the strongest drivers of performance, whereas prompt engineering interventions show diminishing or negative returns in well-equipped environments. Reported performance reflects both model reasoning ability and compatibility with agent tooling and API integration.
Tyler H. Merves, Michael H. Conaway, Joseph M. Escobar +2
Department of Cybersecurity · Department of Information Sciences and Technology
The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against industry standards. We introduceCyberCertBench, a new suite of Multiple Choice Question Answering (MCQA) benchmarks derived from industry recognized certifications. CyberCertBench evaluates LLM domain knowledgeagainst the professional standards of Information Technology cybersecurity and more specializedareas such as Operational Technology and related cybersecurity standards. Concurrently, we propose and validate a novel Proposer-Verifier framework, a methodology to generate interpretable,natural language explanations for model performance. Our evaluation shows that frontier modelsachieve human expert level in general networking and IT security knowledge. However, theiraccuracy declines in questions that require vendor-specific nuances or knowledge in formalstandards, like, e.g., IEC 62443. Analysis of model scaling trend and release date demonstratesremarkable gains in parameter efficiency, while recent larger models show diminishing returns.Code and evaluation scripts are available at: https://github.com/GKeppler/CyberCertBench.
Gustav Keppler, Ghada Elbez, Veit Hagenmeyer
Institute for Automation and Applied Informatics (IAI), Karlsruhe Institute of Technology (KIT), Karlsruhe, Germany
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, domain-specific models and training datasets. To address this gap, we present CyberPal 2.0, a family of cybersecurity-expert small language models (SLMs) ranging from 4B-20B parameters. To train CyberPal 2.0, we generate an enriched chain-of-thought cybersecurity instruction dataset built with our data enrichment and formatting pipeline, SecKnowledge 2.0, which integrates expert-in-the-loop steering of reasoning formats alongside LLM-driven multi-step grounding, yielding higher-fidelity, task-grounded reasoning traces for security tasks. Across diverse cybersecurity benchmarks, CyberPal 2.0 consistently outperforms its baselines and matches or surpasses various open and closed-source frontier models, while remaining a fraction of their size. On core cyber threat intelligence knowledge tasks, our models outperform almost all tested frontier models, ranking second only to Sec-Gemini v1. On core threat-investigation tasks, such as correlating vulnerabilities and bug tickets with weaknesses, our best 20B-parameter model outperforms GPT-4o, o1, o3-mini, and Sec-Gemini v1, ranking first, while our smallest 4B-parameter model ranks second.