LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
Figures & tables
Figure 1 : Why KaliBench Matters? KaliBench enables fine-grained, verifiable evaluation and runtime-free reward training for translating analyst intent into executable Kali commands.
Figure 2 : KaliBench end-to-end pipeline for NL-to-CLI evaluation on Kali Linux. The benchmark is constructed via LLM-based data generation, multi-round verification-regeneration, and representative-sample selection. Evaluation applies alias-aware parsing and fine-grained scoring over tools, arguments, and dimensional categories.
Figure 3 : KaliBench Structure and Representative Sample. KaliBench is a structured benchmark of Kali Linux tools organized along the security attack lifecycle. It comprises 5 security phases and 23 functional dimensions. The dataset sample includes natural-language queries and decomposed ground-truth commands. See Appendix D for a detailed view of examples of all 23 dimensions.
Figure 4 : KaliBench Training Protocol.
Model
Size
Exact Correct (%)
Tool Acc. (%)
Optional F1 (%)
Positional F1 (%)
Total Score (%)
(B)
U
R
H
U
R
H
U
R
H
U
R
H
U
R
H
Avg.
Compact Models ( ≤ 8B)
Foundation-Sec-Ins ⋆
8
14.7
19.1
56.2
66.5
92.5
86.8
33.5
37.4
75.6
53.0
60.0
81.3
51.0
63.3
81.2
65.2
Llama3.1-Ins
8
12.5
18.2
62.7
61.7
93.7
95.4
32.4
37.2
82.7
53.5
61.9
85.8
49.2
64.3
88.0
67.2
Llama-Primus-Base ⋆
8
15.8
19.3
60.3
67.3
90.7
94.1
36.6
39.2
80.8
56.3
61.9
84.0
53.4
63.9
86.3
67.9
Foundation-Sec-Rsn ⋆ †
8
13.3
18.6
66.6
67.0
92.8
95.9
34.8
39.1
85.5
53.0
57.9
86.2
51.6
63.3
89.2
68.0
Table 1 : Open-weight model query-to-command performance across three modes. Columns report Unrestricted (U), Restricted (R), and Hinted (H). † indicates thinking-enabled models, and ⋆ denotes cybersecurity-tuned models. Bold / underline mark the best and second-best within each size group. Models are sorted by Averaged Total Score. MoE parameters use Total/Active (e.g., 685/37B).
Model
Answered queries
EC (%)
EC (%) (all 5K)
GPT-5.6-Sol
4,988
61.83
61.68
Claude Opus 5
3,675
59.89
44.02
Codex (GPT-5.5)
4,633
55.77
51.68
Table 2 : Proprietary models and scaffolded inference in Unrestricted mode.
Figure 5 : Performance across security dimensions and modes. (a) Performance over 23 dimensions and 5 phases per mode. Box plots show distributions (Q1–Q3), with whiskers for min/max. (b,c) Total Score and Exact Correct (%) across modes; dotted lines indicate inter-mode gaps.
Query style
Ures.
Res.
Hinted
Paraphrase
+0.61
+0.22
-0.09
Informal
-0.55
-0.48
-0.28
Messy
-3.20
-1.40
-1.07
Table 3 : Average change in Total Score, in percentage points, relative to original queries across eight model configurations.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Overview of KaliBench samples. Representative examples spanning all five security phases and 23 evaluation dimensions are shown. Zoom in for finer details.
Model Name
Size
Inference
Hugging Face Repository
Llama3.1-Instruct
8B
local vLLM
meta-llama/Llama-3.1-8B-Instruct
Foundation-Sec-Instruct
8B
local vLLM
fdtn-ai/Foundation-Sec-8B
Foundation-Sec-Reasoning
8B
local vLLM
fdtn-ai/Foundation-Sec-8B-Reasoning
Llama-Primus-Base
8B
local vLLM
trend-cybertron/Llama-Primus-Base
Llama-Primus-Merged
8B
local vLLM
trendmicro-ailab/Llama-Primus-Merged
RedSage-Qwen3-8B-Ins
8B
local vLLM
RISys-Lab/RedSage-Qwen3-8B-Ins
Appendix
Table 4: Evaluated open-weight model configurations, parameter counts, inference platforms, and model sources. The inference platform used for each configuration is specified in the table.
Model
Size
Tool score Δ (pp)
Optional F1 Δ (pp)
Positional F1 Δ (pp)
Total score Δ (pp)
Exact Δ (pp)
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Δ1
Δ2
Δ3
Compact Models ( ≤ 8B)
Foundation-Sec-Ins ⋆
8
26.00
-5.70
20.30
3.90
38.20
42.10
7.00
21.30
28.30
12.30
17.90
30.20
4.40
37.10
41.50
Llama3.1-Ins
8
32.00
1.70
33.70
4.80
45.50
50.30
8.40
23.90
32.30
15.10
23.70
38.80
5.70
44.50
50.20
Llama-Primus-Base ⋆
8
23.40
3.40
26.80
2.60
41.60
44.20
5.60
22.10
27.70
10.50
22.40
32.90
3.50
41.00
44.50
Foundation-Sec-Rsn ⋆ †
8
25.80
3.10
28.90
4.30
46.40
50.70
4.90
28.30
33.20
11.70
25.90
37.60
5.30
48.00
53.30
Appendix
Table 5 : Performance deltas across execution modes. We report differences between three execution modes: Unrestricted (U), Restricted (R), and Hinted (H). Δ1 denotes (R − U), Δ2 denotes (H − R), and Δ3 denotes (H − U), all measured in percentage points (pp). † indicates thinking-enabled models, and ⋆ denotes cybersecurity-tuned models. Bold / underline mark the largest and second-largest deltas within each size group. Model order follows Table 1 .
Figure 7 : Coverage of CLI options in KaliBench (Evaluation Subset). While common flags (e.g., --help , --verbose ) appear frequently, many functional tool-specific options also occur across the dataset.
Mode
RedSage-Ins
RedSage-K (GRPO)
ΔGRPO
RedSage-K (SFT+GRPO)
ΔSFT+GRPO
Unrestricted
56.36
67.40
+11.04
70.31
+13.95
Restricted
69.31
74.65
+5.34
77.93
+8.62
Hinted
89.53
88.77
-0.76
89.34
-0.19
Appendix
Table 6 : Effect of RLVR on RedSage-K across evaluation modes. Scores are Total Score percentages. Δ denotes the change from the original RedSage-Ins model in percentage points.
title
avg_total
avg_tool
avg_optional
avg_positional
avg_exact
count
score_sum
Best Reps
d2j-std-apk
1.000
1.000
1.000
1.000
1.000
72
5.000
rfdump
1.000
1.000
1.000
1.000
1.000
72
5.000
truecrypt2john
1.000
1.000
1.000
1.000
1.000
72
5.000
jsql
0.998
1.000
0.995
1.000
0.995
216
4.989
reconspider
0.998
1.000
1.000
0.993
0.993
288
4.984
Appendix
Table 7 : Best- and worst-representative tools aggregated from the scoring files of all models.
Model / Scaffold
Answered queries
Blocked / Missing
TA
Opt-F1
Pos-F1
TS
EC Answered
EC All 5K
GPT-5.6-Sol
4,988 (99.76%)
12 (0.24%)
92.54
87.42
71.51
83.83
61.83
61.68
Claude Opus 5
3,675 (73.50%)
1,325 (26.50%)
90.67
85.95
68.57
81.73
59.89
44.02
Codex (GPT-5.5, xhigh )
4,633 (92.66%)
367 (7.34%)
91.28
82.39
63.46
79.04
55.77
51.68
Appendix
Table 8 : Complete results for proprietary models and agentic scaffolds in the Unrestricted setting. We used the following abbreviations: Tool Accuracy (TA), Optional-Argument F1 (Opt-F1), Positional-Argument F1 (Pos-F1), Total Score (TS), and Exact Correct (EC). All metric scores are percentages. Bold shows the best.
Figure 8 : Model-wise Sensitivity to Messy Queries