cs.CLOct 1, 2026

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Authors: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

Organizations: Khalifa University · University of Western Australia

Abstract

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

    Apr 18, 2026Tyler H. Merves, Michael H. Conaway, Joseph M. Escobar +2Language Model Evasion AttacksLarge Language Model Agents

  2. CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

    Apr 22, 2026Gustav Keppler, Ghada Elbez, Veit HagenmeyerCybersecurityDomain-Specific Language

  3. Toward Cybersecurity-Expert Small Language Models

    Oct 15, 2025Matan Levi, Daniel Ohayon, Ariel Blobstein +3CybersecurityLarge Language Model Safety