Hidden Risks of Jev: An Empirical Study of Security, Privacy, and Dual Use
Authors: Shang Wang, Tianqing Zhu, Huajie Chen, Jiayang Li, Meng Yang, Bo Liu
Organizations: School of Computer Science, University of Technology Sydney, Sydney, Australia · Faculty of Data Science, City University of Macau, Macau, China
Jev turns natural-language questions into typed answers and probabilities with low latency and cost, enabling applications to route requests and select tools. While this interface allows Jev to integrate naturally into application workflows as a decision layer, the security and privacy implications of this emerging use remain largely unexplored. To address this gap, we conduct the first systematic study of these implications using the official Jev API and NanoJev, a local model with controllable training data and updates, focusing on three research questions: (1) What security threats arise when Jev is deployed as an application decision layer? (2) What private information can Jev reveal despite returning constrained typed outputs? (3) How can Jev's general-purpose decision capability be used for beneficial purposes or misused? Jev's decisions depend on application state and may be influenced by user-provided inputs. We therefore adapt prompt injection and adversarial suffixes to manipulate its decisions. Open-source Jev distribution and updates introduce supply-chain risks, which we examine by implanting backdoors in NanoJev through training data poisoning. Since Jev's outputs reflect both application state and information learned during training, we further adapt membership, private attribute, and internal knowledge inference attacks to recover sensitive information despite its constrained output format. Finally, Jev can serve as a general-purpose decision oracle for defensive and malicious workflows. We examine this dual use through four detection tasks covering prompt injection, jailbreak inputs, harmful content, and AI-generated text, alongside misuse scenarios involving jailbreak and model extraction. Our empirical evaluation shows that Jev remains vulnerable to the examined security and privacy threats, while its decision capability can support beneficial and malicious uses.
Figures & tables
Fig. 1 : The top panel illustrates Jev as an application decision layer in a ticket routing workflow, where it evaluates a routing question using the provided state and returns a typed answer with associated probabilities to guide the application’s routing action. The bottom panel summarizes Jev’s characteristics and their associated risks and potential uses, motivating our study.
Fig. 2 : Threat model for input manipulation in a Jev-enabled application. An attacker introduces malicious content into the state through prompt injection or an adversarial suffix, aiming to change Jev’s decision and the subsequent application action while the question and allowed answers remain fixed.
Attack Method
Support routing
Subscription cancellation
Issue severity
Acc ↑
ASR ↑
Acc ↑
ASR ↑
Acc ↑
ASR ↑
Jev
Clean
100.0
0.0
100.0
0.0
100.0
0.0
Context ignoring
3.0
96.5
0.0
100.0
5.5
47.0
Fake completion
96.3
2.5
11.3
88.5
31.3
18.8
Role spoofing
5.8
93.5
100.0
0.0
0.0
86.7
TABLE I : Comparison of six input manipulation attacks against Jev and NanoJev across three tasks. We report Acc and ASR as percentages. Clean rows provide reference accuracy and baseline target prediction rates.
Fig. 3 : Threat model for model poisoning in a Jev-enabled application. An attacker poisons training records to backdoor Jev and uploads the compromised model to a public platform. A victim application then downloads and deploys the model. The attack aims to induce a target decision on trigger-carrying inputs while preserving normal decisions on clean inputs.
Method
Poison rate
cf
james bond
Acc All↑
ASR Target↑
Acc Non↑
Acc All↑
ASR Target↑
Acc Non↑
Base
–
84.0
2.0
95.2
84.0
2.0
95.2
SFT Poisoning
5%
92.3
76.1
94.8
92.6
89.0
94.3
10%
90.4
92.9
94.2
90.5
100.0
91.6
RLCD Poisoning
5%
92.5
82.8
93.1
94.0
98.2
92.7
10%
91.7
96.4
93.8
93.2
100.0
92.5
TABLE II : Backdoor attack results on NanoJev under different poisoning methods, poison rates, and triggers. Results are reported as percentages.
Fig. 4 : Threat model for membership and internal knowledge inference against Jev. An attacker controls the state, question, and answer specification submitted to the model. By observing Jev’s outputs and probabilities, the attacker aims to infer whether a candidate record was used in training or reveal hidden facts learned by the model.
Task
Accuracy ↑
F1 ↑
ROC-AUC ↑
TPR@5%FPR ↑
Support routing
64.6
76.9
55.3
1.9
Cancellation
73.3
82.3
74.6
22.4
Issue severity
61.2
72.9
62.3
3.7
TABLE III : Membership inference results on NanoJev. All metrics are reported as percentages.
Model
Accuracy ↑
ROC-AUC ↑
Jev
97.5
99.6
NanoJev
64.2
65.7
TABLE IV : Internal knowledge inference against Jev and NanoJev using 100 factual and 100 fabricated statements. All metrics are percentages.
Fig. 5 : Internal knowledge inference by year. Bars show the number of factual statements classified as true.
Fig. 6 : Threat model for attribute inference in a Jev-enabled application. An attacker crafts queries associating hidden attribute values in application state with allowed answers. The attacker infers these values from Jev’s outputs, while the application’s question and allowed answers remain fixed.
Model
Access
Attribute
Cov. ↑
Acc. ↑
Jev
Label
Gender
100.0
100.0
Age group
100.0
100.0
Heart disease
100.0
100.0
Ethnicity
100.0
100.0
Probability
Gender
100.0
100.0
Age group
100.0
100.0
TABLE V : Private attribute inference against the official Jev and NanoJev under label and probability access. Cov. and Acc. denote Coverage and Accepted Accuracy. All values are percentages; – indicates undefined accuracy at zero coverage.
Fig. 7 : Jev as a protective tool in LLM-based applications. It returns typed judgments and probabilities to detect prompt injection, identify jailbreak attempts in user inputs and harmful content in model outputs, and distinguish AI-generated from human-written text.
Task
Benchmark
Test (+/–)
Calibration (+/–)
Prompt injection
Tensor Trust
500/500
50/50
BIPIA
180/180
20/20
Jailbreak input
WildJailbreak
180/180
10/10
JailbreakBench
90/90
10/10
Harmful content
JailbreakBench
100/100
20/20
HarmBench
180/180
20/20
TABLE VI : Numbers of positive and negative examples in the test and calibration sets for each benchmark.
Fig. 8 : Detection performance of Jev and NanoJev on four detection tasks.
University of Bonn, Bonn, Germany · Lamarr-Institute for Machine Learning and Artificial Intelligence, Bonn, Germany · Fraunhofer IAIS, Sankt Augustin, Germany