LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation
Organizations: State Key Laboratory of Integrated Services Networks, School of Cyber Engineering, Xidian University · Hangzhou Institute of Technology, Xidian University · Department of Computer Science, Purdue University
Abstract
Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often require substantial ground truth, suffer from low training efficiency, lack support for basic trust properties, and provide limited explainability. Large Language Models (LLMs) offer a compelling alternative due to their strong zero-/few-shot reasoning abilities and broad knowledge. To this end, we propose LLM4Trust, the first benchmark framework that systematically explores the capabilities of LLMs for trust evaluation. We first construct diverse trust graphs to model five basic trust properties and design corresponding property understanding tasks. We then assess the ability of eight representative LLMs to understand these properties under nine prompt methods. Based on this exploration, we identify the most effective LLM-prompt combinations and apply them to five real-world datasets for validating LLMs' trust evaluation capability. During this process, we propose two strategies to extract key information from large-scale trust graphs, addressing the context window limitations of LLMs. Extensive experiments show that LLMs can effectively understand basic trust properties and have great potential for real-world trust evaluation, particularly under limited supervision. However, they remain vulnerable to attacks targeting trust graphs and demonstration examples used in few-shot prompting, and incur high inference costs. Accordingly, we propose a defense mechanism and batch inference to improve the robustness and efficiency of LLM-based trust evaluation. The source code of LLM4Trust is available at https://github.com/Jieerbobo/LLM4Trust
Figures & tables
| Model API | Base Model | # Params | Context Size | Knowledge Cutoff |
| gpt-3.5-turbo | GPT-3.5 | – | 16,385 | 09/2021 |
| gpt-4o | GPT-4o | – | 128k | 10/2023 |
| deepseek-chat | DeepSeek-V3 | 671B | 128k | 07/2024 |
| qwen-max | Qwen-2.5-Max | – | 32k | 10/2023 |
| llama-4-scout-17b-16e-instruct | Llama-4-Scout | 109B | 10M | 08/2024 |
| llama-4-maverick-17b-128e-instruct | Llama-4-Maverick | 400B | 1M | 08/2024 |
| Dataset | # Nodes | # Edges | Avg. Degree | # Trust Levels | Trust Ratio | Graph Type | Domain |
| Advogato | 5,280 | 54,382 | 20.6 | 4 | 32.4% | static | social |
| PGP | 37,841 | 317,081 | 16.8 | 4 | 30.3% | static | security |
| Epinions | 18,098 | 711,508 | 78.6 | 2 | 50.0% | static | e-commerce |
| Bitcoin-OTC | 5,881 | 35,592 | 12.1 | 2 | 90.0% | dynamic | financial |
| Bitcoin-Alpha | 3,783 | 24,186 | 12.8 | 2 | 93.7% | dynamic | financial |
| Task | Model | 0-shot | 1-shot | Few-shot | Know | Role | CoT | CKR | FKR | FCKR |
| Dynamicity | Random | 0.20 | ||||||||
| GPT-3.5 | 0.75 | 0.75 | 0.66 | 0.78 | 0.74 | 0.72 | 0.70 | 0.77 | 0.63 | |
| GPT-4o | 0.85 | 0.96 | 0.96 | 0.84 | 0.81 | 0.85 | 0.83 | 0.98 | 0.98 | |
| DeepSeek-V3 | 0.82 | 0.91 | 0.95 | 0.80 | 0.79 | 0.85 | 0.81 | 0.92 | 0.94 | |
| Qwen-2.5-Max | 0.84 | 0.95 | 0.95 | 0.85 | 0.84 | 0.88 | 0.82 | 0.95 | 0.93 | |
| Llama-4-Scout | 0.82 | 0.92 | 0.86 | 0.81 | 0.85 | 0.82 | 0.81 | 0.90 | 0.91 | |
| Category | Method | Advogato | PGP | Epinions | |||
| F1-micro | MAE | F1-micro | MAE | F1-micro | MAE | ||
| Non- learning | MoleTrust [ 41 ] | 0.584 | 0.309 | 0.640 | 0.332 | 0.691 | 0.364 |
| OpinionWalk [ 36 ] | 0.633 | 0.232 | 0.668 | 0.251 | 0.761 | 0.269 | |
| Fully Supervised | Matri [ 37 ] | 0.650 | 0.141 | 0.673 | 0.136 | 0.776 | 0.150 |
| NeuralWalk [ 38 ] | 0.740 | 0.082 | Out of memory | Out of memory | |||
| Guardian [ 11 ] | 0.730 | 0.087 | 0.870 | 0.084 | 0.879 | 0.097 | |
| Category | Method | Bitcoin-OTC | Bitcoin-Alpha | ||
| F1-macro | BA | F1-macro | BA | ||
| Fully Supervised | Guardian [ 11 ] | 0.613 | 0.613 | 0.530 | 0.540 |
| GATrust [ 39 ] | 0.581 | 0.574 | 0.512 | 0.521 | |
| TrustGNN [ 40 ] | 0.624 | 0.627 | 0.538 | 0.535 | |
| Medley [ 29 ] | 0.622 | 0.595 | 0.571 | 0.558 | |
| DTrust [ 42 ] | 0.637 | 0.658 | 0.583 | 0.607 | |
| Method | Bitcoin-OTC | Bitcoin-Alpha | ||||
| 30% Mix | 30% Time | 3-Shot | 30% Mix | 30% Time | 3-Shot | |
| DeepSeek-V3 | 5.53% | 23.18% | 8.01% | 3.27% | 14.72% | 29.08% |
| Llama-4-Maverick | 2.55% | 12.19% | 1.76% | 1.54% | 9.65% | 10.91% |
| Claude-3.7-Sonnet | 2.72% | 14.04% | 1.81% | 3.91% | 9.50% | 11.02% |
| Dataset | Metric | Batch size | ||||
| 1 | 5 | 10 | 20 | 40 | ||
| PGP | F1-micro | 0.850 | 0.794 | 0.799 | 0.778 | 0.775 |
| Time (s) | 4.821 | 2.825 | 3.860 | 6.940 | 11.960 | |
| Epinions | F1-micro | 0.881 | 0.879 | 0.871 | 0.874 | 0.851 |
| Time (s) | 9.984 | 2.855 | 4.160 | 6.480 | 11.720 | |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Knowledge |
| Dynamicity | Trust is dynamic over time; specifically, is established earlier than if . |
| Asymmetry | Trust is inherently asymmetric; specifically, the trust of node in node , represented as , does not imply an equivalent trust level of node in node , which would be represented as . |
| Conditional Transitivity | Trust is propagative; specifically, the trust of node in node along a given path is determined by the minimum trust level among all edges constituting that path. |
| Composability | Trust is composable; specifically, the trust of node in node along a single path is determined by the minimum trust level among all edges constituting that path. When there are multiple paths from node to node , the trust level from to falls within the range defined by the individual path trust levels. |
| Context-awareness | Trust is context-aware; specifically, means that node trusts node with a trust level of in the context of , but it does not imply that has the same trust in in the context of , which would be represented as . |
| Prompt | Template |
| 0-shot | System <Graph Instruction><Task Instruction><Answer Instruction> User <Question Input> |
| 1-shot | System <Graph Instruction><Task Instruction><Answer Instruction><Example> User <Question Input> |
| Few-shot | System <Graph Instruction><Task Instruction><Answer Instruction><Example 1><Example 2><Example 3> User <Question Input> |
| Knowledge | System <Graph Instruction><Task Instruction><Answer Instruction><Knowledge> User <Question Input> |
| Role | System <Role Definition><Graph Instruction><Task Instruction><Answer Instruction> User <Question Input> |
| CoT | System <Graph Instruction><Task Instruction><Answer Instruction> User <Question Input><CoT> |
| Graph Generator | Small | Medium | Large |
| ER Model | 0.93 | 0.87 | 0.69 |
| SB Model | 0.95 | 0.83 | 0.71 |
| FF Model | 0.91 | 0.94 | 0.77 |
| Dataset | Method | F1-micro | MAE |
| Advogato | Proposed | 0.783 | 0.070 |
| w/o Structural | 0.750 | 0.077 | |
| PGP | Proposed | 0.867 | 0.087 |
| w/o Structural | 0.750 | 0.160 | |
| Epinions | Proposed | 0.950 | 0.040 |
| w/o Structural | 0.650 | 0.280 |
| Dataset | Method | F1-macro | BA |
| Bitcoin-OTC | Proposed | 0.728 | 0.733 |
| w/o Structural | 0.509 | 0.567 | |
| w/o Temporal | 0.620 | 0.633 | |
| Bitcoin-Alpha | Proposed | 0.588 | 0.600 |
| w/o Structural | 0.531 | 0.567 | |
| w/o Temporal | 0.484 | 0.517 |
| Category | Method | Advogato | PGP | Epinions | |||
| Std. | 95% CI | Std. | 95% CI | Std. | 95% CI | ||
| Fully Supervised | Guardian [ 11 ] | 0.006 | [0.723, 0.737] | 0.002 | [0.868, 0.872] | 0.001 | [0.878, 0.881] |
| GATrust [ 39 ] | 0.004 | [0.726, 0.737] | 0.001 | [0.869, 0.872] | 0.001 | [0.879, 0.881] | |
| TrustGNN [ 40 ] | 0.006 | [0.737, 0.751] | 0.001 | [0.880, 0.882] | 0.001 | [0.879, 0.882] | |
| 3-shot GNNs | Guardian [ 11 ] | 0.101 | [0.320, 0.572] | 0.035 | [0.608, 0.694] | 0.041 | [0.732, 0.834] |
| GATrust [ 39 ] | 0.060 | [0.444, 0.592] | 0.066 | [0.632, 0.795] | 0.169 | [0.413, 0.832] | |
| Category | Method | Bitcoin-OTC | Bitcoin-Alpha | ||
| Std. | 95% CI | Std. | 95% CI | ||
| Fully Supervised | Medley [ 29 ] | 0.013 | [0.607, 0.638] | 0.004 | [0.567, 0.576] |
| DTrust [ 42 ] | 0.051 | [0.573, 0.700] | 0.026 | [0.550, 0.615] | |
| TrustGuard [ 10 ] | 0.007 | [0.673, 0.690] | 0.004 | [0.584, 0.595] | |
| 3-shot GNNs | Medley [ 29 ] | 0.022 | [0.468, 0.523] | 0.020 | [0.467, 0.518] |
| DTrust [ 42 ] | 0.044 | [0.381, 0.490] | 0.016 | [0.457, 0.497] | |
| Setting | Advogato | PGP | Epinions | |||
| F1-micro | MAE | F1-micro | MAE | F1-micro | MAE | |
| Original | 0.669 | 0.101 | 0.850 | 0.097 | 0.881 | 0.095 |
| Dataset exposure | 0.662 | 0.102 | 0.851 | 0.097 | 0.885 | 0.092 |
| Node-ID rand. | 0.669 | 0.103 | 0.849 | 0.097 | 0.888 | 0.090 |
| Method | Advogato | PGP | Epinions | Bitcoin -OTC | Bitcoin -Alpha | Total |
| DeepSeek-V3 | 0.63 | 0.21 | 0.35 | 0.45 | 0.40 | 2.04 |
| Llama-4-Maverick | 0.55 | 0.31 | 0.39 | 0.75 | 0.68 | 2.68 |
| Claude-3.7-Sonnet | 7.88 | 4.31 | 5.01 | 8.66 | 8.08 | 33.94 |
| Time | Method | Advogato | PGP | Epinions | Bitcoin -OTC | Bitcoin -Alpha |
| Training (s) | GATrust [ 39 ] | 13.83 | 92.77 | 169.84 | 3.97 | 2.62 |
| TrustGNN [ 40 ] | 60.70 | 519.52 | 502.64 | 5.27 | 3.35 | |
| Medley [ 29 ] | – | – | – | 787.75 | 536.45 | |
| DTrust [ 42 ] | – | – | – | 36.50 | 25.55 | |
| TrustGuard [ 10 ] | – | – | – | 21.90 | 5.65 | |
| Inference (s) | GATrust [ 39 ] | 0.06 | 0.38 | 0.74 | 0.04 | 0.03 |