JevVibe: Efficient Classification-Guided Secure Code Generation
Organizations: Technical University of Darmstadt, Germany
Abstract
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking an autoregressive language model to generate a CWE label and extracting it from the response raises questions about output validity, speed, and cost, as well as accuracy. We evaluate Jev, a decision model that instead selects directly from a declared set of candidates and returns a probability for each, against six open-weight autoregressive models and a frontier proprietary model, GPT-5.6-Sol, on a controlled 50-way CWE classification task over 1,916 CyberSecEval benchmark examples. Jev outperforms all six open-weight baselines on every classification and ranking metric, while its comparison with GPT-5.6-Sol depends on the metric: GPT-5.6-Sol achieves higher Top-1 accuracy and Macro-F1, whereas Jev achieves higher Top-3 and Top-5 accuracy and a nearly identical MRR, at lower median API latency and lower estimated API cost. We further build JevVibe, a diagnosis-guided repair agent that uses predicted CWE labels to repair code generated by Qwen2.5-Coder-32B-Instruct. With Jev providing the diagnosis, the agent increases the detector-measured security pass rate from 63.5% before repair to 70.7%, compared with 66.1% for LLM-guided repair. These results show that JevVibe is effective at improving the security of generated code, with Jev providing reliable and efficient CWE classification.
Figures & tables
| Model | Top-1 | Top-3 | Top-5 | Macro-F1 | MRR | Coverage |
| DeepSeek-Coder-6.7B-Instruct | 0.034 | 0.136 | 0.184 | 0.015 | 0.090 | 0.8205 |
| CodeLlama-34B-Instruct | 0.175 | 0.231 | 0.241 | 0.071 | 0.202 | 0.7114 |
| Llama-3.1-8B-Instruct | 0.177 | 0.274 | 0.306 | 0.084 | 0.228 | 0.9990 |
| DeepSeek-Coder-33B-Instruct | 0.238 | 0.321 | 0.334 | 0.114 | 0.279 | 0.9468 |
| Qwen2.5-Coder-7B-Instruct | 0.303 | 0.468 | 0.532 | 0.131 | 0.385 | 1.0000 |
| Qwen2.5-Coder-32B-Instruct | 0.335 | 0.519 | 0.559 | 0.199 | 0.428 | 0.9995 |
| Metric | GPT-5.6-Sol | Jev |
| Estimated API cost | $15.83 | $0.283 |
| Median latency (s) | 1.531 | 0.244 |
| P95 latency (s) | 3.737 | 0.292 |