cs.CRSep 28, 2026

JevVibe: Efficient Classification-Guided Secure Code Generation

Authors: Arshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi

Organizations: Technical University of Darmstadt, Germany

Abstract

Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at 6.27×6.27\times lower median API latency and 55.9×55.9\times lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.

Figures & tables

Explore similar work

CardsList
  1. Toward Secure Code Generation: Bridging Correctness and Security via Task-Adaptive Vulnerability Modeling and Execution-Based Benchmarking

    Jul 2, 2024Jiexin Wang, Liuwen Cao, Xitong Luo +6Code GenerationVulnerable Code

  2. Enhancing Reliability in LLM-Based Secure Code Generation

    May 22, 2026Mohammed F. Kharma, Mohammad Alkhanafseh, Ahmed Sabbah +1Large Language Model SafetyVulnerable Code